# All Pages Source: https://developer.pdf.co/all Complete index of every PDF.co documentation page, grouped by section. A complete directory of every PDF.co documentation page, grouped by section — a crawlable map of the docs. Jump to any topic below. ## API Reference * [Getting Started](/api) * [Get Account Balance Info](/api/account-balance-info) * [AI Invoice Parser](/api/ai-invoice-parser) * [Understanding Sync and Async Modes](/api/async-and-sync-mode) * [Barcodes Generator](/api/barcode/generate) * [Barcodes Overview](/api/barcode/overview) * [Barcodes Reader](/api/barcode/read) * [Excel to CSV](/api/convert-from-excel/csv) * [Excel to HTML](/api/convert-from-excel/html) * [Excel to JSON](/api/convert-from-excel/json) * [Excel to PDF](/api/convert-from-excel/pdf) * [Excel to Text](/api/convert-from-excel/text) * [Excel to XML](/api/convert-from-excel/xml) * [Document Classifier](/api/document-classifier) * [Document Parser Overview](/api/documentparser/overview) * [Parse Document](/api/documentparser/parser) * [List All Templates](/api/documentparser/templates) * [Retrieve Template by ID](/api/documentparser/templates-id) * [Extract Data from Email File](/api/email/decode) * [Extract Email Attachment](/api/email/extract-attachments) * [Send Email with File](/api/email/send) * [File Download](/api/file-download) * [Delete Temporary File](/api/file-upload/delete) * [Generate Pre-signed URL](/api/file-upload/generate-presigned-url) * [Get MD5 Hash of File by URL](/api/file-upload/hash) * [File Upload Overview](/api/file-upload/overview) * [Upload Small File](/api/file-upload/upload) * [Upload File Using Base64](/api/file-upload/upload-base64) * [Upload File via Pre-signed URL](/api/file-upload/upload-presigned-url-put) * [Upload File from URL \[GET\]](/api/file-upload/upload-url-get) * [Upload File from URL \[POST\]](/api/file-upload/upload-url-post) * [PDF Forms Info Reader](/api/forms/info-reader) * [Background & Job Check](/api/job-check) * [Language Support](/api/language-support) * [Merge PDF](/api/merge/pdf) * [Merge Various Document Type](/api/merge/various-files) * [PDF Add](/api/pdf-add) * [Make Text Searchable](/api/pdf-change-text-searchable/searchable) * [Make Text Unsearchable](/api/pdf-change-text-searchable/unsearchable) * [PDF Compress](/api/pdf-compress) * [PDF Delete Pages](/api/pdf-delete-pages) * [Extract Attachment](/api/pdf-extract-attachments) * [PDF Find Text](/api/pdf-find/basic) * [Find Text in Table with AI](/api/pdf-find/table) * [PDF from CSV](/api/pdf-from-document/csv) * [PDF from DOC](/api/pdf-from-document/doc) * [PDF from Email](/api/pdf-from-email) * [PDF from HTML](/api/pdf-from-html/convert) * [PDF from HTML Template](/api/pdf-from-html/convert-from-template) * [Return HTML Template by ID](/api/pdf-from-html/template-id) * [Return All Templates](/api/pdf-from-html/templates) * [PDF from Image](/api/pdf-from-image) * [PDF from URL](/api/pdf-from-url) * [PDF Info Reader](/api/pdf-info-reader) * [Add Password to PDF](/api/pdf-password/add) * [Remove Password from PDF](/api/pdf-password/remove) * [Auto-rotate Pages with AI](/api/pdf-rotate/auto) * [Rotate Selected Pages](/api/pdf-rotate/basic) * [PDF Search and Delete Text](/api/pdf-search-text-and-delete) * [Search and Replace with Image](/api/pdf-search-text-and-replace/image) * [Search and Replace with Text](/api/pdf-search-text-and-replace/text) * [Split PDF](/api/pdf-split/by-pages) * [Split PDF by Text Search](/api/pdf-split/by-text-search-or-barcode) * [PDF to CSV](/api/pdf-to-csv) * [PDF to XLS](/api/pdf-to-excel/xls) * [PDF to XLSX](/api/pdf-to-excel/xlsx) * [PDF to HTML](/api/pdf-to-html) * [PDF to JPG](/api/pdf-to-image/jpg) * [PDF to PNG](/api/pdf-to-image/png) * [PDF to TIFF](/api/pdf-to-image/tiff) * [PDF to WEBP](/api/pdf-to-image/webp) * [PDF to JSON](/api/pdf-to-json/basic) * [PDF to JSON with AI](/api/pdf-to-json/with-ai) * [PDF to Text](/api/pdf-to-text/basic) * [PDF to Text (Simple)](/api/pdf-to-text/simple) * [PDF to XML](/api/pdf-to-xml) * [Postman](/api/postman) * [Profiles](/api/profiles) * [Response Codes](/api/response-codes) * [Credits per API Function](/api/credits-per-api-function) * [URL Input and Request Limits](/api/url-input-and-request-limits) * [Webhook and Callbacks](/api/webhooks) ## Integrations * [PDF.co Integrations](/integrations) * [Airtable](/integrations/airtable) * [Google Apps Script](/integrations/google-apps-script) * [Add Password and Security into PDF](/integrations/make/add-security-to-pdf) * [Add Text and Images To a PDF](/integrations/make/add-text-images-formfields-to-pdf) * [AI Invoice Parser](/integrations/make/ai-invoice-parser) * [Generate a Barcode](/integrations/make/barcode-generate) * [Read a Barcode](/integrations/make/barcode-read) * [Compress and Optimize PDF](/integrations/make/compress-pdf) * [Convert from PDF](/integrations/make/convert-from-pdf) * [Convert into PDF](/integrations/make/convert-to-pdf) * [Create Fillable PDF Form](/integrations/make/create-fillable-pdf-form) * [Document Classifier](/integrations/make/document-classifier) * [Parse a Document](/integrations/make/document-parser) * [Fill a PDF Form](/integrations/make/fill-pdf-form) * [Getting Started with Make](/integrations/make/getting-started) * [Convert HTML to PDF](/integrations/make/html-to-pdf) * [Convert from Images into PDF](/integrations/make/images-to-pdf) * [Integrating File Sources with PDF.co](/integrations/make/input-file-sources) * [Job Check](/integrations/make/job-check) * [Make Webhooks Integration with PDF.co](/integrations/make/make-webhooks) * [Merge a PDF](/integrations/make/merge) * [Get PDF Information](/integrations/make/pdf-info) * [Convert from PDF into Images](/integrations/make/pdf-to-images) * [Make PDF.co API Call](/integrations/make/pdfco-api-call) * [Remove Password and Security from PDF](/integrations/make/remove-security-from-pdf) * [Search and Delete Found Text in PDF](/integrations/make/search-and-delete-text) * [Search and Replace Text in PDF](/integrations/make/search-and-replace-text) * [Search and Replace With Image in PDF](/integrations/make/search-and-replace-with-image) * [Search Text in PDF](/integrations/make/search-text) * [Send Email With Attachments](/integrations/make/send-email-with-attachments) * [Split a PDF](/integrations/make/split-pdf) * [Upload a File](/integrations/make/upload-file) * [Getting Started with Microsoft Power Automate](/integrations/microsoft-power-automate/getting-started) * [Integrating File Sources with PDF.co](/integrations/microsoft-power-automate/input-file-sources) * [Add Text or Images to PDF](/integrations/n8n/add-text-image-to-pdf) * [AI Invoice Parser](/integrations/n8n/ai-invoice-parser) * [Barcode Generator](/integrations/n8n/barcode-generator) * [Barcode Reader](/integrations/n8n/barcode-reader) * [Compress PDF](/integrations/n8n/compress-pdf) * [Convert PDF to Anything](/integrations/n8n/convert-from-pdf) * [Convert Anything to PDF](/integrations/n8n/convert-to-pdf) * [PDF.co API and n8n Integration Guide](/integrations/n8n/custom-api-call) * [Delete PDF Pages](/integrations/n8n/delete-page-in-pdf) * [Fill a PDF Form](/integrations/n8n/fill-a-pdf-form) * [Get PDF Information & Form Fields](/integrations/n8n/get-pdf-information) * [Getting Started with n8n](/integrations/n8n/getting-started) * [Make PDF Searchable/Unsearchable](/integrations/n8n/make-pdf-searchable-or-unsearchable) * [PDF Merging](/integrations/n8n/merge-pdf) * [PDF Security](/integrations/n8n/pdf-add-remove-security) * [Rotate PDF Pages](/integrations/n8n/rotate-pdf) * [Search and Replace/Delete Text](/integrations/n8n/search-and-replace-text-in-pdf) * [Search in PDF](/integrations/n8n/search-in-pdf) * [PDF Splitting](/integrations/n8n/split-pdf) * [Upload File](/integrations/n8n/upload-file) * [URL/HTML to PDF Conversion](/integrations/n8n/url-html-to-pdf) * [Pabbly Connect](/integrations/pabbly-connect) * [Salesforce](/integrations/salesforce) * [Sharepoint](/integrations/sharepoint) * [Add Barcode](/integrations/zapier/add-barcode-to-pdf) * [Add Form Field](/integrations/zapier/add-formfield-to-pdf) * [Add Image](/integrations/zapier/add-image-to-pdf) * [Add Text](/integrations/zapier/add-text-to-pdf) * [AI Invoice Parser](/integrations/zapier/ai-invoice-parser) * [Anything to PDF](/integrations/zapier/anything-to-pdf) * [Barcode Generator](/integrations/zapier/barcode-generator) * [Advanced Barcode Reader](/integrations/zapier/barcode-reader) * [Compress & Optimize](/integrations/zapier/compress) * [Custom API Call](/integrations/zapier/custom-api-call) * [Document Classifier](/integrations/zapier/document-classifier) * [Document Parser](/integrations/zapier/document-parser) * [Send Email With Attachment](/integrations/zapier/email-send) * [PDF Find Table](/integrations/zapier/find-table) * [Getting Started with Zapier](/integrations/zapier/getting-started) * [HTML to PDF](/integrations/zapier/html-to-pdf) * [Integrating File Sources with PDF.co](/integrations/zapier/input-file-sources) * [Merge PDF](/integrations/zapier/merge) * [PDF Filler](/integrations/zapier/pdf-filler) * [Get PDF Information](/integrations/zapier/pdf-info) * [PDF Page Tools](/integrations/zapier/pdf-page-tools) * [PDF Security](/integrations/zapier/pdf-password-and-security) * [Convert Scanned PDF to Searchable PDF](/integrations/zapier/pdf-searchable) * [PDF to Anything](/integrations/zapier/pdf-to-anything) * [Search and Delete Text](/integrations/zapier/search-and-delete-text) * [Search and Replace Text](/integrations/zapier/search-and-replace-text) * [Search and Replace With Image](/integrations/zapier/search-and-replace-with-image) * [Search Text](/integrations/zapier/search-in-pdf) * [Split PDF Into Multiple Files](/integrations/zapier/split-pdf) * [Split PDF Based on Barcode Search](/integrations/zapier/split-pdf-by-barcode) * [Split PDF Based on Text Search](/integrations/zapier/split-pdf-by-text) ## Knowledge Base * [Overview](/knowledgebase) * [Convert](/knowledgebase/convert) * [Create/Edit](/knowledgebase/create-edit) * [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide) * [Extract](/knowledgebase/extract) * [General](/knowledgebase/general) * [Macros for Text: Built-in and Custom Macros for Auto Text Replacement](/knowledgebase/macros-for-text) * [Manage](/knowledgebase/manage) * [Security](/knowledgebase/security) * [SMTP Configuration Guide](/knowledgebase/smtp-guide) * [How to setup SMTP for email via AOL mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-aol-mail) * [How to Set Up SMTP for Email via GMAIL](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-gmail) * [How to Set Up SMTP for Email via GMX](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-gmx) * [How to Set Up SMTP for Email via Hushmail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-hushmail) * [How to Set Up SMTP for Email via iCloud Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-icloud-mail) * [How to Set Up SMTP for Email via Lycos Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-lycos-mail) * [How to Set Up SMTP for Email via Mail.com](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-mail-com) * [How to Set Up SMTP for Email via Office 365 Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-office-365-mail) * [How to Set Up SMTP for Email via Outlook](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-outlook) * [How to Set Up SMTP for Email via Postmark](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-postmark) * [How to Set Up SMTP for Email via Rediffmail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-rediffmail) * [How to Set Up SMTP for Email via SendGrid](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-sendgrid) * [How to Set Up SMTP for Email via Verizon](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-verizon) * [How to Set Up SMTP for Email via Yahoo Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-yahoo-mail) * [How to Set Up SMTP for Email via Zoho Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-zoho) * [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) ## API Tester * [Welcome to PDF.co API Tester](/api-tester) * [Get Account Balance Info (API Tester)](/api-tester/account-balance-info) * [AI Invoice Parser (API Tester)](/api-tester/ai-invoice-parser) * [Barcodes Generator (API Tester)](/api-tester/barcode/generate) * [Barcodes Reader (API Tester)](/api-tester/barcode/read) * [Excel to CSV (API Tester)](/api-tester/convert-from-excel/csv) * [Excel to HTML (API Tester)](/api-tester/convert-from-excel/html) * [Excel to JSON (API Tester)](/api-tester/convert-from-excel/json) * [Excel to PDF (API Tester)](/api-tester/convert-from-excel/pdf) * [Excel to Text (API Tester)](/api-tester/convert-from-excel/text) * [Excel to XML (API Tester)](/api-tester/convert-from-excel/xml) * [Document Classifier (API Tester)](/api-tester/document-classifier) * [Parse Document (API Tester)](/api-tester/documentparser) * [List All Templates (API Tester)](/api-tester/documentparser/templates) * [Retrieve Template by ID (API Tester)](/api-tester/documentparser/templates-id) * [Extract Data from Email File (API Tester)](/api-tester/email/decode) * [Extract Email Attachment (API Tester)](/api-tester/email/extract-attachments) * [Send Email with File (API Tester)](/api-tester/email/send) * [Delete Temporary File (API Tester)](/api-tester/file-upload/delete) * [Generate Pre-signed URL (API Tester)](/api-tester/file-upload/generate-presigned-url) * [Get MD5 Hash of File by URL (API Tester)](/api-tester/file-upload/hash) * [Upload Small File (API Tester)](/api-tester/file-upload/upload) * [Upload File Using Base64 (API Tester)](/api-tester/file-upload/upload-base64) * [Upload File from URL (API Tester)](/api-tester/file-upload/upload-url-get) * [Upload File from URL (API Tester)](/api-tester/file-upload/upload-url-post) * [PDF Forms Info Reader (API Tester)](/api-tester/forms/info-reader) * [Background & Job Check (API Tester)](/api-tester/job-check) * [Merge PDF (API Tester)](/api-tester/merge/pdf) * [Merge Various Document Type (API Tester)](/api-tester/merge/various-files) * [PDF Add (API Tester)](/api-tester/pdf-add) * [Make Text Searchable (API Tester)](/api-tester/pdf-change-text-searchable/searchable) * [Make Text Unsearchable (API Tester)](/api-tester/pdf-change-text-searchable/unsearchable) * [PDF Compress (API Tester)](/api-tester/pdf-compress) * [PDF Delete Pages (API Tester)](/api-tester/pdf-delete-pages) * [Extract Attachment (API Tester)](/api-tester/pdf-extract-attachments) * [PDF Find Text (API Tester)](/api-tester/pdf-find/basic) * [Find Text in Table with AI (API Tester)](/api-tester/pdf-find/table) * [PDF from CSV (API Tester)](/api-tester/pdf-from-document/csv) * [PDF from DOC (API Tester)](/api-tester/pdf-from-document/doc) * [PDF from Email (API Tester)](/api-tester/pdf-from-email) * [PDF from HTML (API Tester)](/api-tester/pdf-from-html/convert) * [Return HTML Template by ID (API Tester)](/api-tester/pdf-from-html/template-id) * [Return All Templates (API Tester)](/api-tester/pdf-from-html/templates) * [PDF from Image (API Tester)](/api-tester/pdf-from-image) * [PDF from URL (API Tester)](/api-tester/pdf-from-url) * [PDF Info Reader (API Tester)](/api-tester/pdf-info-reader) * [Add Password to PDF (API Tester)](/api-tester/pdf-password/add) * [Remove Password from PDF (API Tester)](/api-tester/pdf-password/remove) * [Auto-rotate Pages with AI (API Tester)](/api-tester/pdf-rotate/auto) * [Rotate Selected Pages (API Tester)](/api-tester/pdf-rotate/basic) * [PDF Search and Delete Text (API Tester)](/api-tester/pdf-search-text-and-delete) * [Search and Replace with Image (API Tester)](/api-tester/pdf-search-text-and-replace/image) * [Search and Replace with Text (API Tester)](/api-tester/pdf-search-text-and-replace/text) * [Split PDF (API Tester)](/api-tester/pdf-split/by-pages) * [Split PDF by Text Search (API Tester)](/api-tester/pdf-split/by-text-search-or-barcode) * [PDF to CSV (API Tester)](/api-tester/pdf-to-csv) * [PDF to XLS (API Tester)](/api-tester/pdf-to-excel/xls) * [PDF to XLSX (API Tester)](/api-tester/pdf-to-excel/xlsx) * [PDF to HTML (API Tester)](/api-tester/pdf-to-html) * [PDF to JPG (API Tester)](/api-tester/pdf-to-image/jpg) * [PDF to PNG (API Tester)](/api-tester/pdf-to-image/png) * [PDF to TIFF (API Tester)](/api-tester/pdf-to-image/tiff) * [PDF to WEBP (API Tester)](/api-tester/pdf-to-image/webp) * [PDF to JSON (API Tester)](/api-tester/pdf-to-json/basic) * [PDF to JSON with AI (API Tester)](/api-tester/pdf-to-json/with-ai) * [PDF to Text (API Tester)](/api-tester/pdf-to-text/basic) * [PDF to Text (Simple) (API Tester)](/api-tester/pdf-to-text/simple) * [PDF to XML (API Tester)](/api-tester/pdf-to-xml) ## Resources * [Changelog](/changelog) # Get Account Balance Info (API Tester) Source: https://developer.pdf.co/api-tester/account-balance-info /openapi.json get /v1/account/credit/balance Run Get Account Balance Info live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Get Account Balance Info → API Reference](/api/account-balance-info) — all parameters, response fields, and limits. # AI Invoice Parser (API Tester) Source: https://developer.pdf.co/api-tester/ai-invoice-parser /openapi.json post /v1/ai-invoice-parser Run AI Invoice Parser live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [AI Invoice Parser → API Reference](/api/ai-invoice-parser) — all parameters, response fields, and limits. ## Prerequisites Before using the AI Invoice Parser API, please note: * **Invoices only**: The API processes invoices exclusively to ensure accurate parsing. * **Asynchronous processing**: When you make a request, you get a JobID immediately while processing happens in the background. To get your results, you can either: * Poll the [**job/check**](/api-tester/job-check) endpoint using your `JobID`, or * Provide a callback URL to get results automatically via webhook. # Barcodes Generator (API Tester) Source: https://developer.pdf.co/api-tester/barcode/generate /openapi.json post /v1/barcode/generate Run Barcodes Generator live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Barcodes Generator → API Reference](/api/barcode/generate) — all parameters, response fields, and limits. # Barcodes Reader (API Tester) Source: https://developer.pdf.co/api-tester/barcode/read /openapi.json post /v1/barcode/read/from/url Run Barcodes Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Barcodes Reader → API Reference](/api/barcode/read) — all parameters, response fields, and limits. # Excel to CSV (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/csv /openapi.json post /v1/xls/convert/to/csv Run Excel to CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to CSV → API Reference](/api/convert-from-excel/csv) — all parameters, response fields, and limits. # Excel to HTML (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/html /openapi.json post /v1/xls/convert/to/html Run Excel to HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to HTML → API Reference](/api/convert-from-excel/html) — all parameters, response fields, and limits. # Excel to JSON (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/json /openapi.json post /v1/xls/convert/to/json Run Excel to JSON live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to JSON → API Reference](/api/convert-from-excel/json) — all parameters, response fields, and limits. # Excel to PDF (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/pdf /openapi.json post /v1/xls/convert/to/pdf Run Excel to PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to PDF → API Reference](/api/convert-from-excel/pdf) — all parameters, response fields, and limits. # Excel to Text (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/text /openapi.json post /v1/xls/convert/to/txt Run Excel to Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to Text → API Reference](/api/convert-from-excel/text) — all parameters, response fields, and limits. # Excel to XML (API Tester) Source: https://developer.pdf.co/api-tester/convert-from-excel/xml /openapi.json post /v1/xls/convert/to/xml Run Excel to XML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Excel to XML → API Reference](/api/convert-from-excel/xml) — all parameters, response fields, and limits. # Document Classifier (API Tester) Source: https://developer.pdf.co/api-tester/document-classifier /openapi.json post /v1/pdf/classifier Run Document Classifier live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Document Classifier → API Reference](/api/document-classifier) — all parameters, response fields, and limits. # Parse Document (API Tester) Source: https://developer.pdf.co/api-tester/documentparser /openapi.json post /v1/pdf/documentparser Run Parse Document live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Parse Document → API Reference](/api/documentparser/overview) — all parameters, response fields, and limits. # List All Templates (API Tester) Source: https://developer.pdf.co/api-tester/documentparser/templates /openapi.json get /v1/pdf/documentparser/templates Run List All Templates live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [List All Templates → API Reference](/api/documentparser/templates) — all parameters, response fields, and limits. # Retrieve Template by ID (API Tester) Source: https://developer.pdf.co/api-tester/documentparser/templates-id /openapi.json get /v1/pdf/documentparser/templates/{id} Run Retrieve Template by ID live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Retrieve Template by ID → API Reference](/api/documentparser/templates-id) — all parameters, response fields, and limits. # Extract Data from Email File (API Tester) Source: https://developer.pdf.co/api-tester/email/decode /openapi.json post /v1/email/decode Run Extract Data from Email File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Extract Data from Email File → API Reference](/api/email/decode) — all parameters, response fields, and limits. # Extract Email Attachment (API Tester) Source: https://developer.pdf.co/api-tester/email/extract-attachments /openapi.json post /v1/email/extract-attachments Run Extract Email Attachment live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Extract Email Attachment → API Reference](/api/email/extract-attachments) — all parameters, response fields, and limits. # Send Email with File (API Tester) Source: https://developer.pdf.co/api-tester/email/send /openapi.json post /v1/email/send Run Send Email with File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Send Email with File → API Reference](/api/email/send) — all parameters, response fields, and limits. # Delete Temporary File (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/delete /openapi.json post /v1/file/delete Run Delete Temporary File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Delete Temporary File → API Reference](/api/file-upload/delete) — all parameters, response fields, and limits. # Generate Pre-signed URL (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/generate-presigned-url /openapi.json get /v1/file/upload/get-presigned-url Run Generate Pre-signed URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Generate Pre-signed URL → API Reference](/api/file-upload/generate-presigned-url) — all parameters, response fields, and limits. # Get MD5 Hash of File by URL (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/hash /openapi.json post /v1/file/hash Run Get MD5 Hash of File by URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Get MD5 Hash of File by URL → API Reference](/api/file-upload/hash) — all parameters, response fields, and limits. # Upload Small File (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/upload /openapi.json post /v1/file/upload Run Upload Small File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Upload Small File → API Reference](/api/file-upload/upload) — all parameters, response fields, and limits. # Upload File Using Base64 (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/upload-base64 /openapi.json post /v1/file/upload/base64 Run Upload File Using Base64 live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Upload File Using Base64 → API Reference](/api/file-upload/upload-base64) — all parameters, response fields, and limits. # Upload File from URL (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/upload-url-get /openapi.json get /v1/file/upload/url Run Upload File from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Upload File from URL → API Reference](/api/file-upload/upload-url-get) — all parameters, response fields, and limits. # Upload File from URL (API Tester) Source: https://developer.pdf.co/api-tester/file-upload/upload-url-post /openapi.json post /v1/file/upload/url Run Upload File from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Upload File from URL → API Reference](/api/file-upload/upload-url-post) — all parameters, response fields, and limits. # PDF Forms Info Reader (API Tester) Source: https://developer.pdf.co/api-tester/forms/info-reader /openapi.json post /v1/pdf/info/fields Run PDF Forms Info Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Forms Info Reader → API Reference](/api/forms/info-reader) — all parameters, response fields, and limits. # Welcome to PDF.co API Tester Source: https://developer.pdf.co/api-tester/index The API Tester helps you quickly explore and test our API endpoints. You can send real requests, check responses, and estimate credit usage, all in one place. It's the fastest way to understand how our API works before you integrate it into your own app or workflow. ## Getting Started To begin testing, you'll need an API Key. * If you already have an account, [log in](https://app.pdf.co/login) to retrieve your key. * If not, [sign up](https://app.pdf.co/signup) and receive 10,000 free credits to start exploring the endpoints. For detailed instructions and examples, see the [Authentication Guide](/api#authenticating-your-api-request). ## URL Input Our API supports files from any public link, such as Google Drive, Dropbox, and others. However, third-party storage providers may restrict the number of requests to their files, which can cause errors during processing. To prevent this, we recommend using our File Upload endpoint to store files in [PDF.co's built-in storage](https://app.pdf.co/tools/files). For details, see our [File Upload guide](/api/file-upload/overview). ## **Response Codes** When you run a request, the tester will return both the **output data** and the **response code**. Common codes include: | Error Code | Description | | :--------- | :------------------------------------------------------------------------------------------------------------------------------- | | `200` | Success. | | `400` | Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password). | | `401` | Unauthorized. Authentication is required and has failed or has not yet been provided. | | `402` | Not enough credits. | | `403` | Access forbidden for input URL. | | `404` | The requested resource could not be found. | [**See Full List of Response Codes**](/api/response-codes) ## Credit Usage Every API call costs credits. The amount depends on the specific endpoint you're using and the size of your file. See [**Credits per API Function**](/api/credits-per-api-function) for the per-endpoint cost, or use our [**Credits Calculator**](https://app.pdf.co/subscriptions#credits-calculator) to estimate usage before sending requests. ## Working with Async Mode :bulb:Tip: For larger or long-running tasks (over 30 seconds), use async mode to avoid timeouts and optimize credit usage. Set the `async` parameter to true when making your request. The API will return a `JobID` and an empty URL, you can then use the [Job Check endpoint](api-tester/job-check) to retrieve the results once processing is complete. Learn more in the [**Sync and Async Mode Guide**](/api/async-and-sync-mode). ## Test Your Endpoint You can try out some of the most popular endpoints right here. Simply choose an endpoint, click **Try It**, and provide the file URL you want to process. Process invoices faster than ever by extracting data and structuring it automatically with our advanced AI. Get quick and accurate data from any invoice, no matter the layout. Add text, images, forms, other PDFs, fill forms, links to external sites and external PDF files. You can update or modify PDF and scanned PDF files. Compress PDF files to reduce their size. Convert PDF and scanned images to text with layout preserved. This method uses OCR and reporoduces layout. # Background & Job Check (API Tester) Source: https://developer.pdf.co/api-tester/job-check /openapi.json post /v1/job/check Run Background & Job Check live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Background & Job Check → API Reference](/api/job-check) — all parameters, response fields, and limits. # Merge PDF (API Tester) Source: https://developer.pdf.co/api-tester/merge/pdf /openapi.json post /v1/pdf/merge Run Merge PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Merge PDF → API Reference](/api/merge/pdf) — all parameters, response fields, and limits. # Merge Various Document Type (API Tester) Source: https://developer.pdf.co/api-tester/merge/various-files /openapi.json post /v1/pdf/merge2 Run Merge Various Document Type live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Merge Various Document Type → API Reference](/api/merge/various-files) — all parameters, response fields, and limits. # PDF Add (API Tester) Source: https://developer.pdf.co/api-tester/pdf-add /openapi.json post /v1/pdf/edit/add Run PDF Add live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Add → API Reference](/api/pdf-add) — all parameters, response fields, and limits. # Make Text Searchable (API Tester) Source: https://developer.pdf.co/api-tester/pdf-change-text-searchable/searchable /openapi.json post /v1/pdf/makesearchable Run Make Text Searchable live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Make Text Searchable → API Reference](/api/pdf-change-text-searchable/searchable) — all parameters, response fields, and limits. # Make Text Unsearchable (API Tester) Source: https://developer.pdf.co/api-tester/pdf-change-text-searchable/unsearchable /openapi.json post /v1/pdf/makeunsearchable Run Make Text Unsearchable live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Make Text Unsearchable → API Reference](/api/pdf-change-text-searchable/unsearchable) — all parameters, response fields, and limits. # PDF Compress (API Tester) Source: https://developer.pdf.co/api-tester/pdf-compress /openapi.json post /v2/pdf/compress Run PDF Compress live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Compress → API Reference](/api/pdf-compress) — all parameters, response fields, and limits. # PDF Delete Pages (API Tester) Source: https://developer.pdf.co/api-tester/pdf-delete-pages /openapi.json post /v1/pdf/edit/delete-pages Run PDF Delete Pages live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Delete Pages → API Reference](/api/pdf-delete-pages) — all parameters, response fields, and limits. # Extract Attachment (API Tester) Source: https://developer.pdf.co/api-tester/pdf-extract-attachments /openapi.json post /v1/pdf/attachments/extract Run Extract Attachment live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Extract Attachment → API Reference](/api/pdf-extract-attachments) — all parameters, response fields, and limits. # PDF Find Text (API Tester) Source: https://developer.pdf.co/api-tester/pdf-find/basic /openapi.json post /v1/pdf/find Run PDF Find Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Find Text → API Reference](/api/pdf-find/basic) — all parameters, response fields, and limits. # Find Text in Table with AI (API Tester) Source: https://developer.pdf.co/api-tester/pdf-find/table /openapi.json post /v1/pdf/find/table Run Find Text in Table with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Find Text in Table with AI → API Reference](/api/pdf-find/table) — all parameters, response fields, and limits. # PDF from CSV (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-document/csv /openapi.json post /v1/pdf/convert/from/csv Run PDF from CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from CSV → API Reference](/api/pdf-from-document/csv) — all parameters, response fields, and limits. # PDF from DOC (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-document/doc /openapi.json post /v1/pdf/convert/from/doc Run PDF from DOC live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from DOC → API Reference](/api/pdf-from-document/doc) — all parameters, response fields, and limits. # PDF from Email (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-email /openapi.json post /v1/pdf/convert/from/email Run PDF from Email live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from Email → API Reference](/api/pdf-from-email) — all parameters, response fields, and limits. # PDF from HTML (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-html/convert /openapi.json post /v1/pdf/convert/from/html Run PDF from HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from HTML → API Reference](/api/pdf-from-html/convert) — all parameters, response fields, and limits. # Return HTML Template by ID (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-html/template-id /openapi.json get /v1/templates/html/{id} Run Return HTML Template by ID live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Return HTML Template by ID → API Reference](/api/pdf-from-html/template-id) — all parameters, response fields, and limits. # Return All Templates (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-html/templates /openapi.json get /v1/templates/html Run Return All Templates live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Return All Templates → API Reference](/api/pdf-from-html/templates) — all parameters, response fields, and limits. # PDF from Image (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-image /openapi.json post /v1/pdf/convert/from/image Run PDF from Image live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from Image → API Reference](/api/pdf-from-image) — all parameters, response fields, and limits. # PDF from URL (API Tester) Source: https://developer.pdf.co/api-tester/pdf-from-url /openapi.json post /v1/pdf/convert/from/url Run PDF from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF from URL → API Reference](/api/pdf-from-url) — all parameters, response fields, and limits. # PDF Info Reader (API Tester) Source: https://developer.pdf.co/api-tester/pdf-info-reader /openapi.json post /v1/pdf/info Run PDF Info Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Info Reader → API Reference](/api/pdf-info-reader) — all parameters, response fields, and limits. # Add Password to PDF (API Tester) Source: https://developer.pdf.co/api-tester/pdf-password/add /openapi.json post /v1/pdf/security/add Run Add Password to PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Add Password to PDF → API Reference](/api/pdf-password/add) — all parameters, response fields, and limits. # Remove Password from PDF (API Tester) Source: https://developer.pdf.co/api-tester/pdf-password/remove /openapi.json post /v1/pdf/security/remove Run Remove Password from PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Remove Password from PDF → API Reference](/api/pdf-password/remove) — all parameters, response fields, and limits. # Auto-rotate Pages with AI (API Tester) Source: https://developer.pdf.co/api-tester/pdf-rotate/auto /openapi.json post /v1/pdf/edit/rotate/auto Run Auto-rotate Pages with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Auto-rotate Pages with AI → API Reference](/api/pdf-rotate/auto) — all parameters, response fields, and limits. # Rotate Selected Pages (API Tester) Source: https://developer.pdf.co/api-tester/pdf-rotate/basic /openapi.json post /v1/pdf/edit/rotate Run Rotate Selected Pages live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Rotate Selected Pages → API Reference](/api/pdf-rotate/basic) — all parameters, response fields, and limits. # PDF Search and Delete Text (API Tester) Source: https://developer.pdf.co/api-tester/pdf-search-text-and-delete /openapi.json post /v1/pdf/edit/delete-text Run PDF Search and Delete Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF Search and Delete Text → API Reference](/api/pdf-search-text-and-delete) — all parameters, response fields, and limits. # Search and Replace with Image (API Tester) Source: https://developer.pdf.co/api-tester/pdf-search-text-and-replace/image /openapi.json post /v1/pdf/edit/replace-text-with-image Run Search and Replace with Image live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Search and Replace with Image → API Reference](/api/pdf-search-text-and-replace/image) — all parameters, response fields, and limits. # Search and Replace with Text (API Tester) Source: https://developer.pdf.co/api-tester/pdf-search-text-and-replace/text /openapi.json post /v1/pdf/edit/replace-text Run Search and Replace with Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Search and Replace with Text → API Reference](/api/pdf-search-text-and-replace/text) — all parameters, response fields, and limits. # Split PDF (API Tester) Source: https://developer.pdf.co/api-tester/pdf-split/by-pages /openapi.json post /v1/pdf/split Run Split PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Split PDF → API Reference](/api/pdf-split/by-pages) — all parameters, response fields, and limits. # Split PDF by Text Search (API Tester) Source: https://developer.pdf.co/api-tester/pdf-split/by-text-search-or-barcode /openapi.json post /v1/pdf/split2 Run Split PDF by Text Search live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [Split PDF by Text Search → API Reference](/api/pdf-split/by-text-search-or-barcode) — all parameters, response fields, and limits. # PDF to CSV (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-csv /openapi.json post /v1/pdf/convert/to/csv Run PDF to CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to CSV → API Reference](/api/pdf-to-csv) — all parameters, response fields, and limits. # PDF to XLS (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-excel/xls /openapi.json post /v1/pdf/convert/to/xls Run PDF to XLS live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to XLS → API Reference](/api/pdf-to-excel/xls) — all parameters, response fields, and limits. # PDF to XLSX (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-excel/xlsx /openapi.json post /v1/pdf/convert/to/xlsx Run PDF to XLSX live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to XLSX → API Reference](/api/pdf-to-excel/xlsx) — all parameters, response fields, and limits. # PDF to HTML (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-html /openapi.json post /v1/pdf/convert/to/html Run PDF to HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to HTML → API Reference](/api/pdf-to-html) — all parameters, response fields, and limits. # PDF to JPG (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-image/jpg /openapi.json post /v1/pdf/convert/to/jpg Run PDF to JPG live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to JPG → API Reference](/api/pdf-to-image/jpg) — all parameters, response fields, and limits. # PDF to PNG (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-image/png /openapi.json post /v1/pdf/convert/to/png Run PDF to PNG live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to PNG → API Reference](/api/pdf-to-image/png) — all parameters, response fields, and limits. # PDF to TIFF (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-image/tiff /openapi.json post /v1/pdf/convert/to/tiff Run PDF to TIFF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to TIFF → API Reference](/api/pdf-to-image/tiff) — all parameters, response fields, and limits. # PDF to WEBP (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-image/webp /openapi.json post /v1/pdf/convert/to/webp Run PDF to WEBP live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to WEBP → API Reference](/api/pdf-to-image/webp) — all parameters, response fields, and limits. # PDF to JSON (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-json/basic /openapi.json post /v1/pdf/convert/to/json2 Run PDF to JSON live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to JSON → API Reference](/api/pdf-to-json/basic) — all parameters, response fields, and limits. # PDF to JSON with AI (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-json/with-ai /openapi.json post /v1/pdf/convert/to/json-meta Run PDF to JSON with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to JSON with AI → API Reference](/api/pdf-to-json/with-ai) — all parameters, response fields, and limits. # PDF to Text (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-text/basic /openapi.json post /v1/pdf/convert/to/text Run PDF to Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to Text → API Reference](/api/pdf-to-text/basic) — all parameters, response fields, and limits. # PDF to Text (Simple) (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-text/simple /openapi.json post /v1/pdf/convert/to/text-simple Run PDF to Text (Simple) live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to Text (Simple) → API Reference](/api/pdf-to-text/simple) — all parameters, response fields, and limits. # PDF to XML (API Tester) Source: https://developer.pdf.co/api-tester/pdf-to-xml /openapi.json post /v1/pdf/convert/to/xml Run PDF to XML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester. **Full reference:** [PDF to XML → API Reference](/api/pdf-to-xml) — all parameters, response fields, and limits. # Get Account Balance Info Source: https://developer.pdf.co/api/account-balance-info Get remaining account balance. **Try it live:** [Get Account Balance Info → API Tester](/api-tester/account-balance-info) — send a real request from your browser. ## `GET /v1/account/credit/balance` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------ | | `remainingCredits` | integer | Number of credits remaining in the account | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/account/credit/balance' \ --header 'x-api-key: ' ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "remainingCredits": 99795868 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/account/credit/balance' \ --header 'x-api-key: ' ``` # AI Invoice Parser Source: https://developer.pdf.co/api/ai-invoice-parser Extract structured data from invoices of any layout using AI-based parsing, without requiring per-vendor templates. **Try it live:** [AI Invoice Parser → API Tester](/api-tester/ai-invoice-parser) — send a real request from your browser. ## `POST /ai-invoice-parser` The AI Invoice parser automatically detects invoice layouts without the manual effort previously required to supply document parsing templates for reference. **Important** * **Only invoices will be parsed**. For all other documents, please use our existing [**Document Parser**](/api/documentparser/parser). * To ensure accurate processing, each invoice must be clearly separated. **If an invoice contains multiple pages, we recommend splitting it** into individual PDFs using the [PDF Split API](/api/pdf-split/by-pages). * While AI Invoice Parser supports multi-page invoices, **the total page count for a single PDF must not exceed 100 pages**. Submitting large PDFs containing multiple invoices is not recommended. * To retrieve results, you must poll the [Background Job Check](/api/job-check) endpoint using the `jobId` returned in the initial response. Once the job status is marked as `success`, the output file will be available at the provided URL. This method extracts data from your PDF invoices and returns a [well-structured JSON format](/api/ai-invoice-parser/#example-response) for your use. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------- | ------ | -------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url attribute`](/api/url-input-and-request-limits#supported-file-sources) | | `customField` | string | *No* | - | JSON string containing [custom field](/api/ai-invoice-parser/#custom-fields) names to extract. Use `camelCase` for field names (e.g., `storeNumber`, `deliveryDate`). Multiple fields should be comma-separated. | | `lineItemStructure` | object | *No* | - | Defines a custom structure for line items in the response. See [Line Item Structure](/api/ai-invoice-parser/#line-item-structure) for more information. | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | | `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | | `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | | `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Custom Fields AI Invoice Parser with custom fields support automatically detects invoice layouts and extracts both standard schema data and user-specified custom fields without requiring manual templates. The `customField` parameter allows you to specify additional fields to extract beyond the standard schema. Some examples include: * `storeNumber` - Store or branch identifier * `deliveryDate` - Expected delivery date * `financialCharges` - Additional financial charges * `lineTotal` - Total amount for line items * `purchaseOrderRef` - Purchase order reference number * `customerReference` - Customer reference number * `departmentCode` - Department or cost center code If a custom field returns an empty value, please [contact our support team](https://pdf.co/support/request?subject=ai-invoice-parser%20-%20custom%20fields) to help improve the extraction accuracy. ## Line Item Structure The `lineItemStructure` attribute lets you define a custom schema for line items. Each key is a field name you choose (in `camelCase`) and each value is the expected data type — either `"string"` or `"number"`. When provided, every object inside the `lineItems` array will contain exactly the fields you specified. If a value cannot be extracted from the invoice, the field will still be present with an empty or default value instead of being omitted. ### Example ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "async": true, "lineItemStructure": { "description": "string", "quantity": "number", "unitPrice": "number", "totalPrice": "number" } } ``` With the structure above, every line item in the response will include all four fields: ```json theme={null} "lineItems": [ [ { "description": "Item 1", "quantity": 2, "unitPrice": 9.95, "totalPrice": 19.90 }, { "description": "Item 2", "quantity": 5, "unitPrice": 20.00, "totalPrice": 100.00 } ] ] ``` Use `camelCase` for field names (e.g., `unitPrice`, `totalPrice`). The field names you define will be used as-is in the response, giving you full control over the output keys. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | -------------------------------------------------------------------------------------- | | `status` | string | Status of the API response. The statuses are: `success`, `error`. | | `message` | string | Descriptive message for the response status. | | `pageCount` | integer | Number of pages processed or returned. | | `body` | object | Contains the invoice data. See [Invoice Schema](#invoice-schema) for more information. | | `jobId` | string | Unique identifier for the background job. | | `credits` | integer | Credits used for this operation. | | `remainingCredits` | integer | Credits left after this job execution. | | `duration` | integer | Time taken to complete the request, in milliseconds. | ### Invoice Schema The `body` object contains all the metadata needed to understand your invoice content and includes the following attributes: ```json theme={null} "body": { "vendor": { .... }, "customer": { .... }, "invoice": { .... }, "paymentDetails": { .... }, "others": { .... }, "lineItems": { .... } } ``` #### Sections * [vendor](#vendor-object) * [customer](#customer-object) * [invoice](#invoice-object) * [paymentDetails](#paymentdetails-object) * [others](#others-object) * [lineItems](#lineitems-object) *** ### The `vendor` Object An `object` containing vendor details. | Attribute | Type | Description | | -------------------- | ------ | ----------------------------------------------------------------------------------- | | `name` | string | Name of the vendor | | `address` | object | Vendor's address details. See [address object](#address-object) | | `contactInformation` | object | Vendor contact details. See [contactInformation object](#contactinformation-object) | | `entityId` | object | Vendor's entity ID (e.g., EIN, ABN, VAT, GST, etc.) | *** ### The `customer` Object An `object` containing customer details. | Attribute | Type | Description | | --------- | ------ | -------------------------------------------------------- | | `billTo` | object | Billing details. See [customer.billTo](#customerbillto) | | `shipTo` | object | Shipping details. See [customer.shipTo](#customershipto) | ### `customer.billTo` | Attribute | Type | Description | | -------------------- | ------ | ----------------------------------------------------------- | | `name` | string | Customer name | | `address` | object | See [address object](#address-object) | | `contactInformation` | object | See [contactInformation object](#contactinformation-object) | | `entityId` | string | Customer's entity ID (e.g., EIN, ABN, VAT, GST, etc.) | ### `customer.shipTo` | Attribute | Type | Description | | --------- | ------ | ------------------------------------- | | `name` | string | Customer name | | `address` | object | See [address object](#address-object) | *** ### The `invoice` Object An `object` containing the invoice details. | Attribute | Type | Description | | ------------- | ------ | --------------------- | | `invoiceNo` | string | Invoice number | | `invoiceDate` | string | Date of invoice | | `poNo` | string | Purchase order number | | `orderNo` | string | Sales order number | *** ### The `paymentDetails` Object An `object` containing payment details. | Attribute | Type | Description | | -------------------- | ------ | ----------------------------------------------------------- | | `paymentTerms` | string | Terms of payment | | `dueDate` | string | Payment due date | | `total` | string | Total amount due | | `subtotal` | string | Subtotal amount | | `tax` | string | Tax amount | | `discount` | string | Discount amount | | `shipping` | string | Shipping amount | | `bankingInformation` | object | See [bankingInformation](#paymentdetailsbankinginformation) | ### `paymentDetails.bankingInformation` | Attribute | Type | Description | | ------------------- | ------ | -------------------------------------------------- | | `bankName` | string | Name of the bank | | `accountHolderName` | string | Name of the account holder | | `accountNumber` | string | Bank account number | | `iban` | string | International Bank Account Number (IBAN) | | `swiftBicCode` | string | SWIFT/BIC code of the bank | | `bankAddress` | object | See [address object](#address-object) | | `bankRoutingCode` | string | Routing code for domestic payments | | `bankCode` | string | Institution number within Canadian banking network | | `branchNumber` | string | Branch-specific code | | `purposeCode` | string | Specifies the transaction's intent | | `additionalNotes` | string | Payment instructions or other notes | *** ### The `others` Object An `object` containing additional notes. | Attribute | Type | Description | | --------- | ------ | ---------------------------------------------- | | `notes` | string | Additional notes such as delivery instructions | *** ### The `lineItems` Object An `object` detailing the line items in an invoice. Note: there is no common structure due to significant variability between invoices! To define your own custom structure, use the [`lineItemStructure`](/api/ai-invoice-parser/#line-item-structure) attribute in your request. A typical invoice might list purchase items with details such as name, quantity or price of each individual item. *** ### Common Objects There are a couble of objects which are commonly used in the schema in a few places, these are as follows. ### `address` object | Attribute | Type | Description | | --------------- | ------ | ----------------- | | `streetAddress` | string | Street address | | `city` | string | City name | | `state` | string | State/county name | | `postalCode` | string | Postal/ZIP code | | `country` | string | Country code/name | ### `contactInformation` object | Attribute | Type | Description | | --------- | ------ | --------------------------- | | `phone` | string | Phone number of the vendor | | `fax` | string | Fax number of the vendor | | `email` | string | Email address of the vendor | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "callback": "https://example.com/callback/url/you/provided" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes). You can use jobId to identify the corresponding callback response. Use the [Job Check](/api/job-check) API to poll the job status. ```json theme={null} { "error": false, "status": "created", "jobId": "7830deca-2e66-11ef-9ad3-8eff830e7461", "credits": 100, "remainingCredits": 106674, "duration": 33 } ``` ## `Example` Callback Response ```json theme={null} { "status": "success", "message": "Success", "pageCount": 1, "body": { "vendor": { "name": "ACME Inc.", "address": { "streetAddress": "1540 Long Street", "city": "Jacksonville", "state": "FL", "postalCode": "32099", "country": "US" }, "contactInformation": { "phone": "352-200-0371", "fax": "904-787-9468" } }, "customer": { "billTo": { "name": "Lanny Lane Ltd.", "address": { "streetAddress": "82 Gorby Lane", "city": "Columbia", "state": "IN", "postalCode": "39429", "country": "US" } }, "shipTo": { "name": "Same as recipient" } }, "invoice": { "invoiceNo": "67893566", "invoiceDate": "JAN 5, 2025" }, "paymentDetails": { "total": "$1,272.35", "subtotal": "$1,262.35", "tax": "$10.00", "shipping": "$0.00" }, "lineItems": [ [ { "quantity": "2", "description": "Item 1", "unit_price": "9.95", "total": "19.90" }, { "quantity": "5", "description": "Item 2", "unit_price": "20.00", "total": "100.00" }, { "quantity": "1", "description": "Item 3", "unit_price": "19.95", "total": "19.95" }, { "quantity": "1", "description": "Item 4", "unit_price": "123.00", "total": "123.00" }, { "quantity": "10", "description": "Item 5", "unit_price": "99.95", "total": "999.50" } ] ] }, "jobId": "7830deca-2e66-11ef-9ad3-8eff830e7461", "credits": 100, "remainingCredits": 106472, "duration": 33 } ``` ## Setting up the Callback URL The callback URL should be a webhook which listens to responses from the parsing results. You can setup your own webhook or use one from a provider. If you are unsure about webhooks or callbacks, please read this [Wikipedia article](https://en.wikipedia.org/wiki/Webhook) to get started. ## Supported Languages * **Albanian (Shqip)** * **Bosnian (Bosanski)** * **Bulgarian (Български)** * **Croatian (Hrvatski)** * **Czech (Čeština)** * **Danish (Dansk)** * **Dutch (Nederlands)** * **English** * **Estonian (Eesti)** * **Finnish (Suomi)** * **French (Français)** * **German (Deutsch)** * **Greek (Ελληνικά)** * **Hungarian (Magyar)** * **Icelandic (Íslenska)** * **Italian (Italiano)** * **Latvian (Latviešu)** * **Lithuanian (Lietuvių)** * **Norwegian (Norsk)** * **Polish (Polski)** * **Portuguese (Português)** * **Romanian (Română)** * **Russian (Русский)** * **Serbian (Српски)** * **Slovak (Slovenčina)** * **Slovenian (Slovenščina)** * **Spanish (Español)** * **Swedish (Svenska)** * **Turkish (Türkçe)** * **Ukrainian (Українська)** **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl -X POST \ https://api.pdf.co/v1/ai-invoice-parser ``` ```javascript theme={null} var https = require("https"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "YOUR_API_KEY_HERE"; // Direct URL of the source PDF file // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"; // Prepare request to `AI Invoice Parser` API endpoint var queryPath = `/v1/ai-invoice-parser`; // JSON payload for api request var jsonPayload = JSON.stringify({ url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; var postRequest = https.request(reqOptions, (response) => { let responseData = ''; response.on("data", (chunk) => { responseData += chunk; }); response.on("end", () => { try { // Parse JSON response var data = JSON.parse(responseData); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log(data.message); } } catch (error) { console.error("Error parsing JSON response:", error); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); function checkIfJobIsCompleted(jobId, resultFileUrl) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { let responseData = ''; response.setEncoding("utf8"); response.on("data", (chunk) => { responseData += chunk; }); response.on("end", () => { try { // Parse JSON response let data = JSON.parse(responseData); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000); } else if (data.status == "success") { console.log("** Response **") console.log(data); } else { console.log(`Operation ended with status: "${data.status}".`); } } catch (error) { console.error("Error parsing JSON response:", error); } }); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf" def main(args = None): getParsedInvoice(SourceFileURL) def getParsedInvoice(uploadedFileUrl): """AI Invoice Parser using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {} parameters["url"] = uploadedFileUrl # Prepare URL for 'AI Invoice Parser' API request url = "{}/ai-invoice-parser".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if(json["status"]): print("** Response **") print(json) return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.Net; using System.Threading; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of Source PDF file // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `AI Invoice Parser` API call string url = "https://api.pdf.co/v1/ai-invoice-parser"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileURL); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Asynchronous job ID string jobId = json["jobId"].ToString(); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { string job_response = ""; string status = CheckJobStatus(jobId, out job_response); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status); if (status == "success") { Console.WriteLine("** Final Response **"); Console.WriteLine(job_response); break; } else if (status == "working") { // Pause for a few seconds Thread.Sleep(3000); } else { Console.WriteLine(status); break; } } while (true); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } static string CheckJobStatus(string jobId, out string response) { using (WebClient webClient = new WebClient()) { // Set API Key webClient.Headers.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; response = webClient.DownloadString(url); JObject json = JObject.Parse(response); return Convert.ToString(json["status"]); } } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.io.OutputStream; import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; import java.time.LocalDateTime; import java.time.format.DateTimeFormatter; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "********************************"; // (!) Make asynchronous job final static boolean Async = true; public static void main(String[] args) throws IOException { // Source PDF file // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"; // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // AI PARSE INVOICE ParseInvoice(webClient, SourceFileUrl); } public static void ParseInvoice(OkHttpClient webClient, String uploadedFileUrl) throws IOException { // Prepare POST request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("url", new JsonPrimitive(uploadedFileUrl)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare URL for AI Invoice Parser API call. // See documentation: https://developer.pdf.co/api/ai-invoice-parser String query = "https://api.pdf.co/v1/ai-invoice-parser"; DateTimeFormatter dtf = DateTimeFormatter.ofPattern("MM/dd/yyyy HH:mm:ss"); // Prepare request to `Document Parser` API Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Asynchronous job ID String jobId = json.get("jobId").getAsString(); System.out.println("Job#" + jobId + ": has been created. - " + dtf.format(LocalDateTime.now())); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { String status = CheckJobStatus(webClient, jobId); // Possible statuses: "working", "failed", "aborted", "success" System.out.println("Job#" + jobId + ": " + status + " - " + dtf.format(LocalDateTime.now())); if (status.compareToIgnoreCase("success") == 0) { break; } else if (status.compareToIgnoreCase("working") == 0) { // Pause for a few seconds try { Thread.sleep(3000); } catch (InterruptedException ex) { Thread.currentThread().interrupt(); // restore interrupted status } } else { System.out.println(status); break; } } while (true); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } // Check Job Status private static String CheckJobStatus(OkHttpClient webClient, String jobId) throws IOException { String url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; String status = ""; // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); status = json.get("status").getAsString(); if(status.equals("success")){ System.out.println(json); } return status; } else { // Display request error System.out.println(response.code() + " " + response.message()); } return "Failed"; } } ``` ```php theme={null} AI Invoice Parser example. "; if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; echo "

== Final Response ==

"; echo $result; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Understanding Sync and Async Modes Source: https://developer.pdf.co/api/async-and-sync-mode Compare Sync and Async request modes, including time limits, jobId-based status checks, and credit cost differences. When you use APIs, the mode you choose Synchronous (Sync) or Asynchronous (Async) can significantly affect both performance and cost. This article explains the differences between these modes, highlights why Async is superior, and shows how switching can benefit you. ## What are Sync and Async Modes? ### Synchronous (Sync) Mode In **Sync** mode, your API request is processed immediately, and the result is returned within a limit of 30 seconds. Here are a few disadvantages: * **API Response**: You receive your results when the processing is complete. * **No Job ID**: There’s no `jobId` provided by [job/check](/api/job-check) for further status checks. * **Time Limit**: If processing takes longer than 30 seconds, your request will fail. * **Higher Costs**: Credits are calculated as: ```javascript theme={null} Total Credits Used = Endpoint Credits × Number of Pages ``` For precise credit calculations, see [Credits per API Function](/api/credits-per-api-function) for the per-endpoint cost, or use the interactive [Credits Calculator](https://app.pdf.co/subscriptions#credits-calculator). ### Asynchronous (Async) Mode To retrieve results, you must poll the [Background Job Check](/api/job-check) endpoint using the `jobId` returned in the initial response. Once the job status is marked as `success`, the output file will be available at the provided URL. **Async** mode processes your request in the background, providing the option to track progress which gives you: * **Immediate API Response**: You receive a `jobId` and an output URL right away. * **Background Processing**: Handles larger tasks without immediate timeouts. * **Extended Time Limit**: Can process tasks for up to 3 minutes, reducing timeouts. * **Status Checks**: Use the `jobId` to monitor progress via the [job/check](/api/job-check) endpoint. * **Webhook Support**: Supports callback to notify you when your job is done on a specified webhook URL. * **Lower Cost**: Credits are calculated as: ```javascript theme={null} Total Credits Used = Endpoint Credits + ("Job/Check" Credits × Number of "Job/Check" Calls until status is "working") ``` ## Why Async Mode is Superior 1. **Handles Bigger Tasks Efficiently**: Designed for larger files or complex tasks that exceed Sync mode’s 30-second limit, reducing failures and saving time. 2. **More Cost-Effective**: Although Async mode includes a small cost for each [job/check](/api/job-check) call (credits charged per check until status is “working”), it often results in overall savings by minimizing failed requests and unnecessary retries. 3. **Better Control and Transparency**: The `jobId` allows you to check your job’s status, giving you more control and clear insight into the process. 4. **Fewer Timeouts and Failures**: Extended time limits decrease the likelihood of failures. 5. **Automates Workflow with Webhooks**: [Webhooks](/api/webhooks) notify you automatically when your job is complete, reducing the need for manual checks and further saving on [job/check](/api/job-check) credits. ## How to Switch to Async Mode 1. **Modify Your API Request**: * Specify that you want to use Async mode in your API call, often by setting an `async` parameter to `true`. 2. **Receive the** `jobId` **and Output URL**: * After submitting your request, you’ll get a `jobId` and an output URL where results will be available once processing is complete. 3. **Implement Status Checks (Optional)**: * Use the [job/check](/api/job-check) endpoint with your `jobId` to monitor progress. * Remember, each [job/check](/api/job-check) call costs credits until status is “working”. 4. **Set Up Webhooks (Recommended)**: * Configure [Webhooks & Callbacks](/api/webhooks) to receive automatic notifications when your job is finished (costs 2 credits). * This reduces the need for manual status checks and saves on [job/check](/api/job-check) credits. ## Conclusion Async mode offers a superior approach for API interactions by providing efficiency, cost-effectiveness, and enhanced control. By switching from Sync to Async mode, you optimize resource usage and unlock features like webhooks and extended processing times. ## Addressing Common Concerns ### What About the Cost of Multiple Status Checks? Async mode is designed to reduce the need for frequent status checks. By setting up [webhooks](/api/webhooks), you receive automatic updates when your job is complete, which minimizes the number of [job/check](/api/job-check) calls and optimizes credit usage. ### Is Switching to Async Mode Complicated? No, it’s straightforward. You handle the `jobId` and output URL provided after your initial request. Implementing [webhooks](/api/webhooks) can make the process even smoother by automating notifications. # Barcodes Generator Source: https://developer.pdf.co/api/barcode/generate Generate high quality barcode images. Supports QR Code, Datamatrix, Code 39, Code 128, PDF417 and many other barcode types. **Try it live:** [Barcodes Generator → API Tester](/api-tester/barcode/generate) — send a real request from your browser. ## `POST /v1/barcode/generate` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `type` | string | *Yes* | QRCode | Set the barcode type to be used. See available barcode types in the [Supported Barcode Types](/api/barcode/overview#supported-barcode-types) | | `value` | string | *Yes* | - | Set the string value to encode inside the barcode, must be in a string format. | | `decorationImage` | string | *No* | - | Set this to the image that you want to be inserted the logo inside the QR-Code barcode. To use your file please upload it first to the temporary storage, see the [Upload Files](/api/file-upload/overview) section. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `Angle` | integer | *No* | `0` | See [profiles.Angle](#profiles-angle) | |     `NarrowBarWidth` | integer | *No* | 3 | See [profiles.NarrowBarWidth](#profiles-narrowbarwidth) | |     `CaptionFont` | string | *No* | Arial, 12 | See [profiles.CaptionFont](#profiles-captionfont) | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | #### `profiles.Angle` Specifies the barcode’s rotation angle as an integer in degrees. | Value | Description | | ----- | --------------------- | | 0 | 0 degrees clockwise | | 1 | 90 degrees clockwise | | 2 | 180 degrees clockwise | | 3 | 270 degrees clockwise | ``` { "profiles": "{'Angle': 3}" } ``` #### `profiles.NarrowBarWidth` Specifies the width of the narrow bars in the barcode in pixels. ``` { "profiles": "{'NarrowBarWidth': 3}" } ``` #### `profiles.CaptionFont` Specifies the font and size of the caption text displayed with the barcode. ``` { "profiles": "{'CaptionFont': 'Arial, 12'}" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ### QRCode Example # ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": false, "async": false, "decorationImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-generator/logo.png" } ``` ### `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/72bc579b37844d9f9e63ce06de5196d8/barcode.png", "error": false, "status": 200, "name": "barcode.png", "duration": 380, "remainingCredits": 98725598, "credits": 7 } ``` #### `Example` CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": false, "async": false }' ``` ### QRCode with Logo Inside Example # ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": true, "async": false } ``` ### `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/9a87556a8b9e4f4eae60843e697250d4/barcode.png", "error": false, "status": 200, "name": "barcode.png", "remainingCredits": 60631 } ``` #### `Example` CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": false, "async": false, "decorationImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-generator/logo.png" }' ``` ### Data URI as Output Example # ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": false, "async": false } ``` ### `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEsAAABLCAYAAAA4TnrqAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAAAOlSURBVHhe7ZBBimMxEMVy/0v34CELkSmBjH96NhaIwKtXlY9fP5fMfawN7mNtcB9rg/tYG9zH2kAf6/V6Pa5hHeaUTPNTjftYg8Z9rEEjPdYJdoc5JZaT0imUOzopywW7w5wSy0npFModnZTlgt1hTonlpHQK5Y5ObJm5SUpODeuU3CSWE53YMnOTlJwa1im5SSwnOrFl5iYpOTWsU3KTWE50YsvMTWI5sY7lxDrMTWI50YktMzeJ5cQ6lhPrMDeJ5UQntszcJJYT61hOrMPcJJYTndgyc5Ps5ob1S24Sy4lObJm5SXZzw/olN4nlRCe2zNwku7lh/ZKbxHKik7JcsDuWE3YosXyXckcnZblgdywn7FBi+S7ljk7KcsHuWE7YocTyXcodnXD5Kck38qc0dDIdOZV8I39KQyfTkVPJN/KnNHzyZaaPrP4v7mNtcB9rA/3n6SOXxHLCDiXTfFmY9j4l03xZ0NZ0cEksJ+xQMs2XhWnvUzLNlwVtTQeXxHLCDiXTfFmY9j4l03xZSK3p+JJYTtgxC9Pe0rAOc2qkr5sOLonlhB2zMO0tDeswp0b6uungklhO2DEL097SsA5zaqSvs0PMi8Zuxzyh3En/YIeYF43djnlCuZP+wQ4xLxq7HfOEcmf7H+yo5WS3Q42puySWk9R5/2bsqOVkt0ONqbsklpPUef9m7KjlZLdDjam7JJaT1Hn/fg1+hElKTo3SIaXfLh3AjzBJyalROqT026UD+BEmKTk1SoeUfrv0BdLHHXSYUyN13r+/Tvq4gw5zaqTO+/fXSR930GFOjdR5//4Dl5+SWF7gLiWWk9Ih2uKhpySWF7hLieWkdIi2eOgpieUF7lJiOSkdoq3dQ8bJHe5SY+oujam7NHRSlgsnd7hLjam7NKbu0tBJWS6c3OEuNabu0pi6S0MntszcJJYb7NPCbp+UXZ3YMnOTWG6wTwu7fVJ2dWLLzE1iucE+Lez2SdnViS0zN0nJTWPqVsk0Xxo6sWXmJim5aUzdKpnmS0MntszcJCU3jalbJdN8aejElpmbxPJdyh12KNnNiU5smblJLN+l3GGHkt2c6MSWmZvE8l3KHXYo2c2JTspywe4wp8TywsmuoZee+jO7w5wSywsnu4ZeeurP7A5zSiwvnOwaeol/9pTEcsIONabup8RyQ1s89JTEcsIONabup8RyQ1s89JTEcsIONabup8Ryo7Uuf7mPtcF9rA3uY21wH2uD+1iZn58/9whzEbhRquEAAAAASUVORK5CYII=", "error": false, "status": 200, "name": "barcode.png", "duration": 298, "remainingCredits": 98725605, "credits": 7 } ``` #### `Example` CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "name": "barcode.png", "value": "abcdef123456", "type": "QRCode", "inline": true, "async": false }' ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```javascript theme={null} var https = require("https"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Result image file name const DestinationFile = "./barcode.png"; // Barcode type. See valid barcode types in the documentation https://developer.pdf.co const BarcodeType = "Code128"; // Barcode value const BarcodeValue = "qweasd123456"; // Prepare request to `Barcode Generator` API endpoint var queryPath = `/v1/barcode/generate`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: 'barcode.png', type: BarcodeType, value: BarcodeValue }); var reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; exports.handler = async (event) => { let dataString = ''; const promise_response = await new Promise((resolve, reject) => { // Send request var postRequest = https.request(reqOptions, (response) => { response.on('data', chunk => { dataString += chunk; }); response.on('end', () => { resolve({ statusCode: 200, body: JSON.stringify(JSON.parse(dataString), null, 4) }); }); }).on("error", (e) => { reject({ statusCode: 500, body: 'Something went wrong!' }); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); }); return promise_response; }; ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Result file name ResultFile = ".\\barcode.png" # Barcode type. See valid barcode types in the documentation https://developer.pdf.co BarcodeType = "Code128" # Barcode value BarcodeValue = "qweasd123456" def main(args = None): generateBarcode(ResultFile) def generateBarcode(destinationFile): """Generates Barcode using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/barcode/generate parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["type"] = BarcodeType parameters["value"] = BarcodeValue # Prepare URL for 'Barcode Generate' API request url = "{}/barcode/generate".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFCOWebApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Result file name const string ResultFileName = @".\barcode.png"; // Barcode type. See valid barcode types in the documentation https://developer.pdf.co const string BarcodeType = "Code128"; // Barcode value const string BarcodeValue = "qweasd123456"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare requests params as JSON // See documentation: https://developer.pdf.co/#barcode-generator Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(ResultFileName)); parameters.Add("type", BarcodeType); parameters.Add("value", BarcodeValue); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // URL of "Barcode Generator" endpoint string url = "https://api.pdf.co/v1/barcode/generate"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated barcode image file string resultFileURI = json["url"].ToString(); // Download generated image file webClient.DownloadFile(resultFileURI, ResultFileName); Console.WriteLine("Generated barcode saved to \"{0}\" file.", ResultFileName); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } finally { webClient.Dispose(); } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Result file name final static Path ResultFile = Paths.get(".\\barcode.png"); // Barcode type. See valid barcode types in the documentation https://developer.pdf.co final static String BarcodeType = "Code128"; // Barcode value final static String BarcodeValue = "qweasd123456"; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Barcode Generator` API call String query = "https://api.pdf.co/v1/barcode/generate"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"type\": \"%s\", \"value\": \"%s\"}", ResultFile.getFileName(), BarcodeType, BarcodeValue); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated barcode image file String resultFileUrl = json.get("url").getAsString(); // Download the image file downloadFile(webClient, resultFileUrl, ResultFile); System.out.printf("Generated barcode saved to \"%s\" file.", ResultFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile.toFile()); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} ## Result:"; } else { // Display service reported errors echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> ```
# Barcodes Overview Source: https://developer.pdf.co/api/barcode/overview Generate and read barcodes. ## Supported Barcode Types | Name | Type | Character Set | Length | Notes | | --------------------------- | ------- | ------------------------------------------------------------------------------------- | ---------------------------------------- | ------------------------------------------------------------------- | | `AustralianPostCode` | 2D | Numbers Only | 4 | - | | `Aztec` | 2D | Full ASCII; FNC1 and ESI control codes | Variable, Min 12 - Max 3832 | - | | `Codabar` | Linear | Numbers: 0-9; Symbols: - : . \$ / + Start/Stop Characters: A, B, C, D, E, \*, N, or T | Variable | - | | `CodablockF` | Complex | - | - | See this guide. | | `Code128` | Linear | All ASCII characters and control codes | Variable | - | | `Code16K` | - | - | - | - | | `Code39` | Linear | Uppercase letters A-Z; Numbers 0-9; Space - . \$ / + % | Variable | - | | `Code39Extended` | Linear | All ASCII characters and control codes | Variable | - | | `Code39Mod43` | - | - | - | - | | `Code39Mod43Extended` | - | - | - | - | | `Code93` | Linear | Uppercase letters A-Z; Numbers 0-9; Space - . \$ / + % | - | - | | `DataMatrix` | 2D | All ASCII characters | Variable | - | | `DPMDataMatrix` | - | - | - | - | | `EAN13` | Linear | Numbers Only | 13 + check digit +2 optional +5 optional | - | | `EAN2` | Linear | Numbers Only | Exact 2 Numbers | - | | `EAN5` | Linear | Numbers Only | Exact 5 Numbers | - | | `EAN8` | Linear | Numbers Only | 7 + check digit | - | | `GS1 - 128` | Linear | ASCII symbols | 128 ASCII symbols | - | | `GS1DataBarExpanded` | Linear | String | 74 numeric or 41 alphabetic characters | - | | `GS1DataBarExpandedStacked` | Linear | String | 74 numeric or 41 alphabetic characters | - | | `GS1DataBarLimited` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified | | `GS1DataBarOmnidirectional` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified | | `GS1DataBarStacked` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified | | `GTIN12` | Linear | Numbers Only | Expects 11 digits; 12th optional | - | | `GTIN13` | Linear | Numbers Only | Expects 12 digits; 13th optional | - | | `GTIN14` | Linear | Numbers Only | Expects 13 digits; 14th optional | - | | `GTIN8` | Linear | Numbers Only | Expects 7 digits; 8th optional | - | | `IntelligentMail` | Linear | Numbers Only | Up to 31 digits | Tracking: 20 digits; Rounding: 0,5,9,11 digits; Spaces/dots allowed | | `Interleaved2of5` | Linear | Numbers Only | - | EVEN if no checksum; ODD if checksum added | | `ITF14` | Linear | Numbers Only | Expects 13 digits; 14th optional | Will be verified | | `MaxiCode` | 2D | All ASCII characters | 93 | - | | `MICR` | - | - | - | - | | `MicroPDF` | 2D | String | 1850 characters or 2710 digits | - | | `MSI` | Linear | Numbers Only | Variable | - | | `PatchCode` | - | - | - | - | | `PDF417` | Linear | All 256 ASCII characters and 8-bit binary data | Variable | - | | `Pharmacode` | Linear | Decimal Numbers | 1 to 131070 | - | | `PostNet` | Postal | Numbers Only | 5 + check digit +4 optional +6 optional | - | | `PZN` | Linear | Numbers Only | Exactly 6 or 7 digits | - | | `QRCode` | 2D | All ASCII characters | Variable | - | | `RoyalMail` | Postal | Digits and characters from A to Z | - | - | | `RoyalMailKIX` | Postal | All numeric digits (0-9), uppercase letters (A-Z) | Variable | - | | `Trioptic` | - | - | - | - | | `UPCA` | Linear | Numbers Only | 11 + check digit +2 optional +5 optional | - | | `UPCE` | Linear | Numbers Only | 8 digits total | Encodes 6 digits + number system + check digit | | `UPU` | Postal | - | - | - | # Barcodes Reader Source: https://developer.pdf.co/api/barcode/read Read barcodes from images and **PDF**. Can read all popular barcode types from QR Code and Code 128, EAN to Datamatrix, PDF417, GS1 and many other barcodes. **Try it live:** [Barcodes Reader → API Tester](/api-tester/barcode/read) — send a real request from your browser. ## `POST /v1/barcode/read/from/url` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `type` | string | *Yes* | QRCode | Set the barcode type to be used. See available barcode types in the [Supported Barcode Types](/api/barcode/overview#supported-barcode-types) | | `types` | string | *No* | - | Detects checkboxes, radiobuttons, vertical and horizontal lines, and general segments (all content types) on scanned documents using the barcode reader engine. Comma-separated list of object types to decode, must be in a string format. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `RenderingResolution` | integer | *No* | 120 | Set the rendering resolution for the barcode reader engine. The default resolution is 120 DPI. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `barcodes` | array\[object] | List of barcodes found in the document | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ### types Detects checkboxes, radiobuttons, vertical and horizontal lines, and general segments (all content types) on scanned documents using the barcode reader engine. Comma-separated list of object types to decode, must be in a string format. ## Visual Element Detection Modes * Checkbox: Locates check boxes. * Segment: Locates and selects objects on a page (general selection). * UnderlinedField: Detects fillable fields (typically, underlined spaces, i.e. fields to fill in a form). * Rectangle: Detects rectangles, including checkboxes. Also returns the value as 1 if a checkmark or a filled rectangle was detected. * Oval: Detects rounded or oval marks (typically, a radiobutton). Returns value of 1 if filled out radiobutton was detected. * HorizontalLine: Detects horizontal lines. * VerticalLine: Detects vertical lines. ## Normal Example ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf", "types": "Checkbox,UnderlinedField", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "barcodes": [ { "Value": "abcdef123456", "RawData": "", "Type": 14, "Rect": "{X=448,Y=23,Width=106,Height=112}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 1, "Metadata": "", "TypeName": "QRCode" }, { "Value": "test123", "RawData": "", "Type": 2, "Rect": "{X=111,Y=60,Width=255,Height=37}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 0.90625155, "Metadata": "", "TypeName": "Code128" }, { "Value": "123456", "RawData": "", "Type": 4, "Rect": "{X=111,Y=129,Width=306,Height=37}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 0.7710818, "Metadata": "", "TypeName": "Code39" }, { "Value": "0112345678901231", "RawData": "", "Type": 2, "Rect": "{X=111,Y=198,Width=305,Height=37}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 0.9156459, "Metadata": "", "TypeName": "Code128" }, { "Value": "12345670", "RawData": [ 1, 2, 3, 4, 5, 6, 7, 0 ], "Type": 5, "Rect": "{X=111,Y=267,Width=182,Height=0}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 1, "Metadata": "", "TypeName": "I2of5" }, { "Value": "1234567890128", "RawData": "", "Type": 6, "Rect": "{X=102,Y=336,Width=71,Height=72}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "Confidence": 0.895925164, "Metadata": "", "TypeName": "EAN13" } ], "pageCount": 1, "error": false, "status": 200, "remainingCredits": 99826192, "credits": 35 } ``` #### `Example` CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/barcode/read/from/url' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "types": "QRCode,Code128,Code39,Interleaved2of5,EAN13", "pages": "0", "async": false }' ``` ## Optical Marks Reader Our barcode reader engine can also find the following marks and objects on scanned documents: * Checkboxes * Radioboxes * Vertical and horizontal lines * General segments (basically, all content types on the page). ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf", "types": "QRCode,Code128,Code39,Interleaved2of5,EAN13", "pages": "0", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "barcodes": [ { "Value": "box", "RawData": "", "Type": 53, "Rect": "{X=298,Y=437,Width=132,Height=6}", "Page": 0, "File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf", "Confidence": 1, "Metadata": "", "TypeName": "UnderlinedField" } ], "pageCount": 1, "error": false, "status": 200, "duration": 860, "remainingCredits": 98725528, "credits": 35 } ``` #### `Example` CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/barcode/read/from/url' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf", "types": "Checkbox,UnderlinedField", "async": false }' ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```javascript theme={null} var https = require("https"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source file to search barcodes in. const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf"; // Comma-separated list of barcode types to search. // See valid barcode types in the documentation https://developer.pdf.co const BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // Prepare request to `Barcode Reader` API endpoint var queryPath = `/v1/barcode/read/from/url`; // JSON payload for api request var jsonPayload = JSON.stringify({ types: BarcodeTypes, pages: Pages, url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Display found barcodes in console data.barcodes.forEach((element) => { console.log("Found barcode:"); console.log(" Type: " + element.TypeName); console.log(" Value: " + element.Value); console.log(" Document Page Index: " + element.Page); console.log(" Rectangle: " + element.Rect); console.log(" Confidence: " + element.Confidence); console.log(""); }, this); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.error(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source file to search barcodes in. SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf" # Comma-separated list of barcode types to search. # See valid barcode types in the documentation https://developer.pdf.co BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" def main(args=None): readBarcodes(SourceFileURL) def readBarcodes(uploadedFileUrl): """Get Barcode Information using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {} parameters["types"] = BarcodeTypes parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'Barcode Reader' API request url = "{}/barcode/read/from/url".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Display information for barcode in json["barcodes"]: print("Found barcode:") print(f" Type: {barcode['TypeName']}") print(f" Value: {barcode['Value']}") print(f" Document Page Index: {barcode['Page']}") print(f" Rectangle: {barcode['Rect']}") print(f" Confidence: {barcode['Confidence']}") print("") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source file (image or PDF) to search barcodes in. const string SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf"; // Comma-separated list of barcode types to search. // See valid barcode types in the documentation https://developer.pdf.co const string BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare requests params as JSON // See documentation: https://developer.pdf.co Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileURL); parameters.Add("type", BarcodeTypes); parameters.Add("pages", Pages); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // URL of "Barcode Reader" endpoint string url = "https://api.pdf.co/v1/barcode/read/from/url"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Display found barcodes in console foreach (JToken token in json["barcodes"]) { Console.WriteLine("Found barcode:"); Console.WriteLine(" Type: " + token["TypeName"]); Console.WriteLine(" Value: " + token["Value"]); Console.WriteLine(" Document Page Index: " + token["Page"]); Console.WriteLine(" Rectangle: " + token["Rect"]); Console.WriteLine(" Confidence: " + token["Confidence"]); Console.WriteLine(); } } else { // Display service reported error Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { // Display request error Console.WriteLine(e.ToString()); } finally { webClient.Dispose(); } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source file to search barcodes in. final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf"; // Comma-separated list of barcode types to search. // See valid barcode types in the documentation https://developer.pdf.co final static String BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Barcode Reader` API call String query = "https://api.pdf.co/v1/barcode/read/from/url"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"types\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", BarcodeTypes, Pages, SourceFileURL); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Display found barcodes in console for (JsonElement element : json.get("barcodes").getAsJsonArray()) { JsonObject barcode = (JsonObject) element; System.out.println("Found barcode:"); System.out.println(" Type: " + barcode.get("TypeName").getAsString()); System.out.println(" Value: " + barcode.get("Value").getAsString()); System.out.println(" Document Page Index: " + barcode.get("Page").getAsString()); System.out.println(" Rectangle: " + barcode.get("Rect").getAsString()); System.out.println(" Confidence: " + barcode.get("Confidence").getAsString()); System.out.println(); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } } ``` ```php theme={null} Cloud API asynchronous "Barcode Reader" job example (allows to avoid timeout errors). "; if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // URL of generated JSON file that will available after the job completion $resultFileUrl = $json["url"]; // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to JSON file with information about decoded barcodes echo "
## Conversion Result:" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Excel to CSV Source: https://developer.pdf.co/api/convert-from-excel/csv Converts a `xls`/`xlsx` file to `csv`. **Try it live:** [Excel to CSV → API Tester](/api-tester/convert-from-excel/csv) — send a real request from your browser. ## `POST /v1/xls/convert/to/csv` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/xls/convert/to/csv' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false }' ``` ```python theme={null} import requests import json # Your API endpoint URL. url = "https://api.pdf.co/v1/xls/convert/to/csv" # Your API Key. api_key = "Your API Key" # The URL of the Excel file you want to convert. input_file_url = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls" headers = { "x-api-key": api_key, "Content-Type": "application/json" } data = { "url": input_file_url, "async": False } response = requests.post(url, headers=headers, json=data) if response.status_code == 200: # The request was successful. # Parse the json response. data = response.json() # Extract the CSV file URL from the response. csv_url = data.get('url', '') print("CSV file is available at: ", csv_url) else: # There was an error with the request. print("Error: ", response.status_code) ``` # Excel to HTML Source: https://developer.pdf.co/api/convert-from-excel/html Converts a `xls`/`xlsx`/`csv` file to `html`. **Try it live:** [Excel to HTML → API Tester](/api-tester/convert-from-excel/html) — send a real request from your browser. ## `POST /v1/xls/convert/to/html` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` # Excel to JSON Source: https://developer.pdf.co/api/convert-from-excel/json Converts a `xls`/`xlsx`/`csv` file to `json`. **Try it live:** [Excel to JSON → API Tester](/api-tester/convert-from-excel/json) — send a real request from your browser. ## `POST /v1/xls/convert/to/json` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` # Excel to PDF Source: https://developer.pdf.co/api/convert-from-excel/pdf Converts a `xls`/`xlsx`/`csv` file to `pdf`. **Try it live:** [Excel to PDF → API Tester](/api-tester/convert-from-excel/pdf) — send a real request from your browser. ## `POST /v1/xls/convert/to/pdf` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination CSV file name const DestinationFile = "./result.csv"; // Prepare request to `PDF To CSV` API endpoint var queryPath = `/v1/xls/convert/to/csv`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download CSV file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated CSV file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests import json # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source xls file. SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls" def main(args = None): convertXlsToXml(SourceFileURL) def convertXlsToXml(sourceFileUrl): """Convert Xls/Xlsx to Xml using PDF.co Web API""" # Prepare requests params as JSON parameters = { "url": sourceFileUrl, "async": False } # Prepare URL for 'Xls to Xml' API request url = "{}/xls/convert/to/xml".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, headers={ "x-api-key": API_KEY, "Content-Type": "application/json" }, data=json.dumps(parameters)) if (response.status_code == 200): json_res = response.json() if json_res["error"] == False: # Get URL of result file resultFileUrl = json_res["url"] # Output URL of converted xml file print(f"Result file url: {resultFileUrl}") else: # Show service reported error print(json_res["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` # Excel to Text Source: https://developer.pdf.co/api/convert-from-excel/text Converts a `xls`/`xlsx`/`csv` file to `text`. **Try it live:** [Excel to Text → API Tester](/api-tester/convert-from-excel/text) — send a real request from your browser. ## `POST /v1/xls/convert/to/txt` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` # Excel to XML Source: https://developer.pdf.co/api/convert-from-excel/xml Converts a `xls`/`xlsx`/`csv` file to `xml`. **Try it live:** [Excel to XML → API Tester](/api-tester/convert-from-excel/xml) — send a real request from your browser. ## `POST /v1/xls/convert/to/xml` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "async": false, "name": "Output" } ``` # Credits per API Function Source: https://developer.pdf.co/api/credits-per-api-function Estimate PDF.co API credit usage by endpoint. See how many credits each API function uses and whether the charge is per page or per API call. Every PDF.co API endpoint consumes credits when you call it. This page lists the credit cost of each endpoint so you can estimate how many credits a workflow will use before you run it. Endpoints are charged in one of two ways, **per processed page** or **per API call**, and the tables below are grouped accordingly. For a quick estimate, use the interactive **Credit Calculator** on the [Subscriptions](https://app.pdf.co/subscriptions) page. It lets you pick an endpoint and page count and shows the credit cost instantly. ## How to estimate credits | Usage type | Formula | | ------------ | ---------------------------------------------------- | | Per page | `total credits = endpoint credits × pages processed` | | Per API call | `total credits = endpoint credits × API calls` | For example, converting 5,000 one-page PDFs to JPG with `/v1/pdf/convert/to/jpg` uses `5,000 × 12 = 60,000 credits`. If each document has multiple pages, multiply by the total number of pages. For example, 5,000 documents with 3 pages each is 15,000 pages, so PDF to JPG would use `15,000 × 12 = 180,000 credits`. ## Endpoints charged per page | Endpoint | API path | Credits (per page) | | ------------------------------------------------ | -------------------------------------- | -----------------: | | Parse invoices with AI | `/v1/ai-invoice-parser` | 100 | | Extract data with Document Parser template | `/v1/pdf/documentparser` | 42 | | Add text, images, and fill form fields | `/v1/pdf/edit/add` | 21 | | Merge PDFs into one | `/v1/pdf/merge` | 2 | | Merge images, documents, and PDFs into a new PDF | `/v1/pdf/merge2` | 35 | | Split PDF by page numbers | `/v1/pdf/split` | 2 | | Split PDF by text search | `/v1/pdf/split2` | 35 | | Convert XLS or XLSX to PDF | `/v1/xls/convert/to/pdf` | 21 | | Convert CSV to PDF | `/v1/pdf/convert/from/csv` | 21 | | Convert DOC, DOCX, RTF, TXT, or XPS to PDF | `/v1/pdf/convert/from/doc` | 21 | | Convert HTML to PDF | `/v1/pdf/convert/from/html` | 9 | | Convert images to PDF | `/v1/pdf/convert/from/image` | 9 | | Convert URL to PDF | `/v1/pdf/convert/from/url` | 9 | | Convert EML or MSG to PDF | `/v1/pdf/convert/from/email` | 56 | | Convert PDF to CSV (AI-powered) | `/v1/pdf/convert/to/csv` | 28 | | Convert PDF to HTML | `/v1/pdf/convert/to/html` | 21 | | Convert PDF to JSON (legacy) | `/v1/pdf/convert/to/json` | 28 | | Convert PDF to JSON (AI-powered) | `/v1/pdf/convert/to/json2` | 28 | | Convert PDF to JSON with metadata (AI-powered) | `/v1/pdf/convert/to/json-meta` | 42 | | Convert PDF to text (AI-powered) | `/v1/pdf/convert/to/text` | 21 | | Convert PDF to text (simple, no AI) | `/v1/pdf/convert/to/text-simple` | 4 | | Convert PDF to XLS (AI-powered) | `/v1/pdf/convert/to/xls` | 35 | | Convert PDF to XLSX (AI-powered) | `/v1/pdf/convert/to/xlsx` | 28 | | Convert PDF to XML (AI-powered) | `/v1/pdf/convert/to/xml` | 35 | | Render PDF to JPG | `/v1/pdf/convert/to/jpg` | 12 | | Render PDF to PNG | `/v1/pdf/convert/to/png` | 15 | | Render PDF to WebP | `/v1/pdf/convert/to/webp` | 18 | | Render PDF to TIFF | `/v1/pdf/convert/to/tiff` | 28 | | Convert XLS or XLSX to CSV | `/v1/xls/convert/to/csv` | 9 | | Convert XLS or XLSX to HTML | `/v1/xls/convert/to/html` | 9 | | Convert XLS or XLSX to JSON | `/v1/xls/convert/to/json` | 15 | | Convert XLS or XLSX to TXT | `/v1/xls/convert/to/txt` | 9 | | Convert XLS or XLSX to XML | `/v1/xls/convert/to/xml` | 15 | | Rotate pages | `/v1/pdf/edit/rotate` | 7 | | Detect and fix page rotation | `/v1/pdf/edit/rotate/auto` | 28 | | Remove pages from PDF | `/v1/pdf/edit/delete-pages` | 5 | | Replace text in PDF | `/v1/pdf/edit/replace-text` | 21 | | Replace text with an image in PDF | `/v1/pdf/edit/replace-text-with-image` | 77 | | Delete text in PDF | `/v1/pdf/edit/delete-text` | 21 | | Read barcodes from URL or file | `/v1/barcode/read/from/url` | 35 | | Extract attachments from MSG or EML | `/v1/email/extract-attachments` | 35 | | Decode email from MSG or EML | `/v1/email/decode` | 35 | | Send email with attachments | `/v1/email/send` | 21 | | Add security protection to PDF | `/v1/pdf/security/add` | 3 | | Remove protection from PDF | `/v1/pdf/security/remove` | 3 | | Extract PDF attachments | `/v1/pdf/attachments/extract` | 8 | | Classify document based on rules | `/v1/pdf/classifier` | 42 | | Compress PDF to reduce file size | `/v2/pdf/compress` | 35 | | Read PDF file information | `/v1/pdf/info` | 7 | | Find text inside PDFs and images | `/v1/pdf/find` | 35 | | Return JSON with table information | `/v1/pdf/find/table` | 21 | | Convert scanned PDF to text-searchable PDF | `/v1/pdf/makesearchable` | 35 | | Convert PDF to scanned (unsearchable) PDF | `/v1/pdf/makeunsearchable` | 35 | ## Endpoints charged per API call | Endpoint | API path | Credits (per API call) | | --------------------------------------------- | -------------------------------------- | ---------------------: | | Generate barcode image | `/v1/barcode/generate` | 7 | | Generate file upload URL | `/v1/file/upload/get-presigned-url` | 7 | | Upload a small local file as a temporary file | `/v1/file/upload` | 11 | | Upload file from URL | `/v1/file/upload/url` | 11 | | Upload file from Base64 | `/v1/file/upload/base64` | 21 | | Get HTML templates | `/v1/templates/html` | 2 | | Get Document Parser templates | `/v1/pdf/documentparser/templates` | 2 | | Get Document Parser template by ID | `/v1/pdf/documentparser/templates/:id` | 2 | | Read PDF form fields | `/v1/pdf/info/fields` | 8 | | Check background job status | `/v1/job/check` | 2 | ## Notes * [Estimated credits](/api/async-and-sync-mode) are calculated for sync mode and could differ in async mode. * Some workflows make more than one API call. For example, uploading a file separately before processing it consumes upload endpoint credits in addition to the processing endpoint credits. * Async jobs may require status checks with `/v1/job/check`, which is listed separately in the per API call table. * To check the remaining credits in an account, use [`/v1/account/credit/balance`](/api/account-balance-info). * You can also estimate costs with the [Credit Calculator](https://app.pdf.co/subscriptions) on the Subscriptions page. # Document Classifier Source: https://developer.pdf.co/api/document-classifier Detect the class of an incoming PDF, JPG, or PNG document using keyword rules or built-in AI, so you can route it to the right processing template. **Try it live:** [Document Classifier → API Tester](/api-tester/document-classifier) — send a real request from your browser. ## `POST /v1/pdf/classifier` Document Classifier can automatically find class of input PDF, JPG, PNG document by analyzing its content using the built-in AI or custom defined classification rules. The best way to **develop**, **test** and **maintain** classification rules is to use `Classifier Tester Tool` from PDF.co [Document Classifier UI](https://app.pdf.co/document-classifier) . Use this tool to quickly edit and test rules on single PDFs and on folders. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. | | `rulescsv` | string | *No* | - | Define custom classification rules in CSV format. See the [rulescsv](#rulescsv). | | `rulescsvurl` | string | *No* | - | URL to the CSV file with classification rules. For the format, see the description above `rulescsv` parameter | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `RenderTextObjects` | boolean | *No* | `true` | Render text objects or not | |     `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not | |     `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not | |     `TIFFCompression` | string | *No* | `LZW` | TIFF compression algorithm. The options are: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE` | |     `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters.AddGammaCorrection()` | array\[string (float format)] | *No* | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. | |     `OCRImagePreprocessingFilters.AddGrayscale()` | boolean | *No* | `false` | Set to true to preprocessing filter that converts a colored document/image to grayscale before performing OCR | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### `rulescsv` Rules are in CSV format where each row contains: `class name`, `logic` (`AND` or `OR` (default)), and keywords separated by a comma. Each row is separated by the `\n` symbol. You can use regular expressions for keywords with this syntax: `/keyword or regexp/i` where `i` is the case-insensitive flag. Please note that all `\` symbols should add the prefix `\` because of JSON format, so `\d` becomes `\\d` and so on. > **Custom Rules Example 1** for `rulescsv`. > > ``` > Amazon AWS, OR, Amazon Web Services Invoice, Amazon CloudFront\nDigital Ocean, OR,DigitalOcean, DOInvoice\nACME,OR, ACME Inc.,1540 Long Street > ``` > **Custom Rules Example 2**. > > ``` > Medical Report,AND,/Instructing Party|Medical Report|Date Of Injury|Med Agency Ref/i\r\nInjured Claimant,OR, Injured Claimant, Injured Patient ID > ``` ## Document Classifier Usage Guide This Document Classifier checks content of input PDF, JPG, PNG, or TIFF. It uses AI to automatically determine the class of the document (e.g., `finance`, `invoice`) and returns the result to the user. Custom-defined classification rules can also be used. Use this Document Classifier to quickly build a workflow for sorting input documents and PDF files. ### How to Create and Test Custom Classification Rules Classification rules are stored in CSV format, one line per class, with the following format: ``` className, logicType, keyword1, keyword2, keyword3 ... ``` Where: * `className` – The name of the class. It will be returned if rules from this class match the document. * `logicType` – (Optional) Logic to use for keywords. Can be `OR` (default) or `AND`. `OR` means the class is identified if one or more keywords match. `AND` means **all** keywords must match. If not specified, `OR` is assumed. * `keyword1`, `keyword2`, `keyword3` – Keywords or phrases to check. Can include regular expressions, e.g., `/\d+/` or `/Medical Report|Med Report/i`. ### Sample Rules ``` Invoice,OR,Invoice Number,Invoice #,Invoice No,Tax Invoice,, Purchase Order,OR,PO Number,Order Number,Order No,,, Bill,OR,Bill Date,Billing Period,Bill Number,,, Bank Statement,OR,/Account Statement/i,/Statement of Account/i,Business Checking,Accounts Payable,/Statement No/i, Income Statement,OR,/Income Statement/i,,,,, Has US Number,OR,"/\b-?(\d+,?)+(\.\d\d)\b/",,,,, Medical Report,AND,/Medical Report|Med Report/i ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `body` | object | Response body. | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes).. For more information, see [Response Codes](/api/response-codes). | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "async": false, "inline": "true", "password": "", "profiles": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "classes": [ { "class": "invoice" }, { "class": "finance" }, { "class": "documents" } ] }, "pageCount": 1, "error": false, "status": 200, "credits": 42, "duration": 353, "remainingCredits": 98019328 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/classifier' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "async": false, "inline": "true", "password": "", "profiles": "" } ' ``` ```javascript theme={null} var request = require('request'); var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/pdf/classifier', 'headers': { 'Content-Type': 'application/json', 'x-api-key': 'YOUR_PDFCO_API_KEY' }, body: JSON.stringify({ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "async": false, "encrypt": "false", "inline": "true", "password": "", "profiles": "" }) }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ``` ```csharp theme={null} using System; using RestSharp; namespace HelloWorldApplication { class HelloWorld { static void Main(string[] args) { var client = new RestClient("https://api.pdf.co/v1/pdf/classifier"); client.Timeout = -1; var request = new RestRequest(Method.POST); request.AddHeader("Content-Type", "application/json"); request.AddHeader("x-api-key", "YOUR_PDFCO_API_KEY"); var body = @"{" + "\n" + @" ""url"": ""https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf""," + "\n" + @" ""async"": false," + "\n" + @" ""encrypt"": ""false""," + "\n" + @" ""inline"": ""true""," + "\n" + @" ""password"": """"," + "\n" + @" ""profiles"": """"" + "\n" + @"} "; request.AddParameter("application/json", body, ParameterType.RequestBody); IRestResponse response = client.Execute(request); Console.WriteLine(response.Content); } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("application/json"); RequestBody body = RequestBody.create(mediaType, "{\n \"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf\",\n \"async\": false,\n \"encrypt\": \"false\",\n \"inline\": \"true\",\n \"password\": \"\",\n \"profiles\": \"\"\n} "); Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/classifier") .method("POST", body) .addHeader("Content-Type", "application/json") .addHeader("x-api-key", "YOUR_PDFCO_API_KEY") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} 'https://api.pdf.co/v1/pdf/classifier', CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => '', CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => 'POST', CURLOPT_POSTFIELDS =>'{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "async": false, "encrypt": "false", "inline": "true", "password": "", "profiles": "" } ', CURLOPT_HTTPHEADER => array( 'Content-Type: application/json', 'x-api-key: YOUR_PDFCO_API_KEY' ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; ?> ```
# Document Parser Overview Source: https://developer.pdf.co/api/documentparser/overview Parse PDFs and scanned documents to extract fields, tables, values, and barcodes from invoices, statements, orders, and similar forms. **Try it live:** [Parse Document → API Tester](/api-tester/documentparser) — send a real request from your browser. ## Built-in document parser templates `General Invoice Template` can parse invoices (English only) to invoice id, invoice date, extract total, tax, and line items. Set the `templateId` parameter to `1` to use this template. ## How to classify incoming documents before parsing them? Use the [/pdf/classifier](/api/document-classifier) endpoint (see below) to automatically sort/detect the class of the document based on AI or on custom keywords-based rules. For example, you can easily define rules to find which vendor provided the document to find which template to apply accordingly. See [Document Classifier](https://developer.pdf.co/api/document-classifier) for more details. ## Additional Information and Tools * [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) * [PDF.co Document Parser: Template Creation Guide](https://developer.pdf.co/knowledgebase/document-parser-guide) # Parse Document Source: https://developer.pdf.co/api/documentparser/parser Extract fields, tables, and values from PDFs and images using Document Parser templates with searchable form fields and multi-page support. ## `POST /v1/pdf/documentparser` Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `templateId` | integer | *No* | - | Set ID of document parser template to be used. View and manage your templates at [Document Parser](https://app.pdf.co/document-parser/templates) | | `template` | string | *No* | - | The raw format of the document parser template to be used directly. see [Template](/api/documentparser/parser) | | `password` | string | *No* | - | Password for the PDF file. | | `inline` | boolean | *No* | `false` | Set to true to include the results directly in the response, in addition to providing a URL to the generated output file. Applies only when `async` mode is enabled. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `outputFormat` | string | *No* | `JSON` | The format of the output file. The output format can be `JSON`, `CSV`, or `XML`. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | | `body` | object | *No* | |     `objects` | array\[object] | | |     `elapsed` | float | Processing time in seconds | |     `templateName` | string | Name of the parsing template used | |     `templateVersion` | string | Version of the parsing template | |     `timestamp` | string | Timestamp when the parsing occurred | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "outputFormat": "JSON", "templateId": "1", "async": false, "inline": "true", "password": "", "profiles": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "objects": [ { "name": "companyName", "objectType": "field", "value": "Amazon Web Services, Inc", "rectangle": [ 0, 0, 0, 0 ] }, { "name": "companyName2", "objectType": "field", "value": "Amazon Web Services, Inc", "rectangle": [ 0, 0, 0, 0 ] }, { "name": "invoiceId", "objectType": "field", "value": "123456789", "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "name": "dateIssued", "objectType": "field", "value": "2018-04-03T00:00:00", "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "name": "dateDue", "objectType": "field", "value": "2018-04-03T00:00:00", "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "name": "bankAccount", "objectType": "field", "value": "123456789012", "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "name": "total", "objectType": "field", "value": 6.58, "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "name": "subTotal", "objectType": "field", "value": "" }, { "name": "tax", "objectType": "field", "value": 1.01, "pageIndex": 0, "rectangle": [ 0, 0, 0, 0 ] }, { "objectType": "table", "name": "table", "rows": [] } ], "templateName": "Generic Invoice [en]", "templateVersion": "4", "timestamp": "2020-08-21T19:23:31" }, "pageCount": 1, "error": false, "status": 200, "name": "sample-invoice.json", "remainingCredits": 60803 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/documentparser' \ --header 'Content-Type: application/json' \ --header 'x-api-key: {{x-api-key}}' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf", "outputFormat": "JSON", "templateId": "1", "async": false, "inline": "true", "password": "", "profiles": "" }' ``` ```javascript theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.io.OutputStream; import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { // Source PDF file final Path SourceFile = Paths.get(".\\MultiPageTable.pdf"); // PDF document password. Leave empty for unprotected documents. final String Password = ""; // Destination JSON file name final Path DestinationFile = Paths.get(".\\result.json"); // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. PARSE UPLOADED PDF DOCUMENT ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String uploadedFileUrl, String templateText) throws IOException { // Prepare POST request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("url", new JsonPrimitive(uploadedFileUrl)); jsonBody.add("template", new JsonPrimitive(templateText)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request to `Document Parser` API Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/documentparser") .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated JSON file String resultFileUrl = json.get("url").getAsString(); // Download JSON file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\MultiPageTable.pdf" # Destination JSON file name DestinationFile = ".\\result.json" // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) # to create templates. # Read template from file: file_read = open(".\\MultiPageTable-template1.yml", mode='r', encoding="utf-8",errors="ignore") Template = file_read.read() file_read.close() def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): PerformDocumentParser(uploadedFileUrl, Template, DestinationFile) def PerformDocumentParser(uploadedFileUrl, template, destinationFile): # Content data = { 'url': uploadedFileUrl, 'template': template } # Prepare URL for 'Document Parser' API request url = "{}/pdf/documentparser".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data= data, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"}) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.IO; using System.Net; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\MultiPageTable.pdf"; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination TXT file name const string DestinationFile = @".\result.json"; static void Main(string[] args) { // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: String templateText = File.ReadAllText(@".\MultiPageTable-template1.yml"); // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. PARSE UPLOADED PDF DOCUMENT // URL of `Document Parser` API call string url = "https://api.pdf.co/v1/pdf/documentparser"; Dictionary requestBody = new Dictionary(); requestBody.Add("template", templateText); requestBody.Add("name", Path.GetFileName(DestinationFile)); requestBody.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(requestBody); // Execute request response = webClient.UploadString(url, "POST", jsonPayload); // Parse response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated JSON file string resultFileUrl = json["url"].ToString(); // Download JSON file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.io.OutputStream; import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { // Source PDF file final Path SourceFile = Paths.get(".\\MultiPageTable.pdf"); // PDF document password. Leave empty for unprotected documents. final String Password = ""; // Destination JSON file name final Path DestinationFile = Paths.get(".\\result.json"); // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. PARSE UPLOADED PDF DOCUMENT ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String uploadedFileUrl, String templateText) throws IOException { // Prepare POST request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("url", new JsonPrimitive(uploadedFileUrl)); jsonBody.add("template", new JsonPrimitive(templateText)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request to `Document Parser` API Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/documentparser") .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated JSON file String resultFileUrl = json.get("url").getAsString(); // Download JSON file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Document Parse Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ParseDocument($apiKey, $uploadedFileUrl, $templateText) { // (!) Make asynchronous job $async = TRUE; // Prepare URL for Document parser API call. // See documentation: https://developer.pdf.co/api/documentparser/parser $url = "https://api.pdf.co/v1/pdf/documentparser"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["template"] = $templateText; $parameters["async"] = $async; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); echo $result . "
"; if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // URL of generated JSON file that will available after the job completion $resultFileUrl = $json["url"]; // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to JSON file with information about parsed fields echo "

Parsing Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# List All Templates Source: https://developer.pdf.co/api/documentparser/templates Returns all **Document Parser** data extraction templates available to the current user. **Try it live:** [List All Templates → API Tester](/api-tester/documentparser/templates) — send a real request from your browser. ## `GET /v1/pdf/documentparser/templates` Use the PDF.co dashbaord to manage your [Document Parser Templates](https://app.pdf.co/document-parser/templates). Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | -------------- | ------------------------------------------ | | `templates` | array\[object] | | | `remainingCredits` | integer | Number of credits remaining in the account | | `credits` | integer | Number of credits consumed by the request | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "templates": [ { "id": 40, "type": "user", "title": "Untitled", "description": "Untitled" }, { "id": 1, "type": "system", "title": "Invoice Parser", "description": "Parses invoices and extracts invoice number, company name, due date, amount, tax" } ], "remainingCredits": 94229 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/pdf/documentparser/templates' \ --header 'Content-Type: application/json' \ --header 'x-api-key: {{x-api-key}}' ``` # Retrieve Template by ID Source: https://developer.pdf.co/api/documentparser/templates-id Returns detailed information for document parser template by template's id. **Try it live:** [Retrieve Template by ID → API Tester](/api-tester/documentparser/templates-id) — send a real request from your browser. ## `GET /v1/pdf/documentparser/templates/:id` Use the PDF.co dashbaord to manage your [Document Parser Templates](https://app.pdf.co/document-parser/templates). Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ----------------- | -------------------------------------------------------------------- | | `id` | integer | Unique identifier for the template | | `type` | string | Source of the template. The available sources are: `user`, `system`. | | `title` | string | Title of the template | | `description` | string | Description of what the template does | | `created_at` | String (ISO 8601) | Timestamp indicating when the template was initially created | | `updated_at` | String (ISO 8601) | Timestamp indicating the last time the template was modified | | `body` | string | Template content | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/pdf/documentparser/templates/1' \ --header 'Content-Type: application/json' \ --header 'x-api-key: {{x-api-key}}' \ --data-raw '' ``` ```javascript theme={null} /*jshint esversion: 6 */ var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./MultiPageTable.pdf"; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.json"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. OPTIMIZE UPLOADED PDF FILE parsePdf(API_KEY, uploadedFileUrl, Password, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function parsePdf(apiKey, uploadedFileUrl, password, destinationFile) { // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: var templateText = fs.readFileSync("./MultiPageTable-template1.yml", "utf-8"); // URL for `Document Parser` API call var query = `https://api.pdf.co/v1/pdf/documentparser`; var jsonRequestObject = { url: uploadedFileUrl, template: templateText }; request( { url: query, headers: { "x-api-key": API_KEY }, method: "POST", json: true, body: jsonRequestObject }, function (error, response, body) { if (error) { return console.error("Error: ", error); } // Parse JSON response let data = JSON.parse(JSON.stringify(body)); if (data.error == false) { //Download generated file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated result file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("Error: " + data.message); } } ); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\MultiPageTable.pdf" # Destination JSON file name DestinationFile = ".\\result.json" // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) # to create templates. # Read template from file: file_read = open(".\\MultiPageTable-template1.yml", mode='r', encoding="utf-8",errors="ignore") Template = file_read.read() file_read.close() def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): PerformDocumentParser(uploadedFileUrl, Template, DestinationFile) def PerformDocumentParser(uploadedFileUrl, template, destinationFile): # Content data = { 'url': uploadedFileUrl, 'template': template } # Prepare URL for 'Document Parser' API request url = "{}/pdf/documentparser".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data= data, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"}) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.IO; using System.Net; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\MultiPageTable.pdf"; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination TXT file name const string DestinationFile = @".\result.json"; static void Main(string[] args) { // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: String templateText = File.ReadAllText(@".\MultiPageTable-template1.yml"); // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. PARSE UPLOADED PDF DOCUMENT // URL of `Document Parser` API call string url = "https://api.pdf.co/v1/pdf/documentparser"; Dictionary requestBody = new Dictionary(); requestBody.Add("template", templateText); requestBody.Add("name", Path.GetFileName(DestinationFile)); requestBody.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(requestBody); // Execute request response = webClient.UploadString(url, "POST", jsonPayload); // Parse response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated JSON file string resultFileUrl = json["url"].ToString(); // Download JSON file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.io.OutputStream; import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { // Source PDF file final Path SourceFile = Paths.get(".\\MultiPageTable.pdf"); // PDF document password. Leave empty for unprotected documents. final String Password = ""; // Destination JSON file name final Path DestinationFile = Paths.get(".\\result.json"); // Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser) // to create templates. // Read template from file: String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. PARSE UPLOADED PDF DOCUMENT ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String uploadedFileUrl, String templateText) throws IOException { // Prepare POST request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("url", new JsonPrimitive(uploadedFileUrl)); jsonBody.add("template", new JsonPrimitive(templateText)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request to `Document Parser` API Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/documentparser") .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated JSON file String resultFileUrl = json.get("url").getAsString(); // Download JSON file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Document Parse Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ParseDocument($apiKey, $uploadedFileUrl, $templateText) { // (!) Make asynchronous job $async = TRUE; // Prepare URL for Document parser API call. // See documentation: https://developer.pdf.co/api/documentparser/parser $url = "https://api.pdf.co/v1/pdf/documentparser"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["template"] = $templateText; $parameters["async"] = $async; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); echo $result . "
"; if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // URL of generated JSON file that will available after the job completion $resultFileUrl = $json["url"]; // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to JSON file with information about parsed fields echo "

Parsing Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Extract Data from Email File Source: https://developer.pdf.co/api/email/decode Decode an email message to extract its components. **Try it live:** [Extract Data from Email File → API Tester](/api-tester/email/decode) — send a real request from your browser. ## `POST /v1/email/decode` For converting email to PDF please see [PDF from Email](/api/pdf-from-email). ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `responseParameters` | object | *No* | - | - | |     `body` | object | *No* | - | Response body. | |     `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) | |     `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | |     `name` | string | *No* | - | Name of the output file | |     `credits` | integer | *No* | - | Number of credits consumed by the request | |     `remainingCredits` | integer | *No* | - | Number of credits remaining in the account | |     `duration` | integer | *No* | - | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml", "inline": true, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "from": "test@example.com", "fromName": "", "to": [ { "address": "test2@example.com", "name": "" } ], "cc": [], "bcc": [], "sentAt": null, "receivedAt": null, "subject": "Test email with attachments", "bodyHtml": null, "bodyText": "Test Email Message with 2 PDF files as attachments\r\n\r\n", "attachmentCount": 2 }, "error": false, "status": 200, "name": "sample.json", "remainingCredits": 60095 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/email/decode' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml", "inline": true, "async": false }' ``` # Extract Email Attachment Source: https://developer.pdf.co/api/email/extract-attachments Extract attachments from an email **Try it live:** [Extract Email Attachment → API Tester](/api-tester/email/extract-attachments) — send a real request from your browser. ## `POST /v1/email/extract-attachments` For converting email to PDF please see [PDF from Email](/api/pdf-from-email). ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `responseParameters` | object | *No* | - | - | |     `body` | object | *No* | - | Response body. | |     `pageCount` | integer | *No* | - | Number of pages in the PDF document. | |     `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) | |     `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | |     `name` | string | *No* | - | Name of the output file | |     `credits` | integer | *No* | - | Number of credits consumed by the request | |     `remainingCredits` | integer | *No* | - | Number of credits remaining in the account | |     `duration` | integer | *No* | - | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml", "inline": true, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "from": "test@example.com", "subject": "Test email with attachments", "bodyHtml": null, "bodyText": "Test Email Message with 2 PDF files as attachments\r\n\r\n", "attachments": [ { "filename": "DigitalOcean.pdf", "url": "https://pdf-temp-files.s3.amazonaws.com/2943e6bb80e646ec92e839292e95d542/DigitalOcean.pdf" }, { "filename": "sample.pdf", "url": "https://pdf-temp-files.s3.amazonaws.com/e10e37fbb438432a83ece50ccdc719b3/sample.pdf" } ] }, "pageCount": 2, "error": false, "status": 200, "name": "sample.json", "remainingCredits": 60085 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/email/extract-attachments' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml", "inline": true, "async": false }' ``` ```javascript theme={null} var request = require('request'); var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/email/send', 'headers': { 'Content-Type': 'application/json', 'x-api-key': 'ADD_YOUR_PDFco_API_KEY' }, body: JSON.stringify({ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "from": "John Doe ", "to": "Partner ", "subject": "Check attached sample pdf", "bodytext": "Please check the attached pdf", "bodyHtml": "Please check the attached pdf", "smtpserver": "smtp.gmail.com", "smtpport": "587", "smtpusername": "my@gmail.com", "smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833", "async": false }) }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ``` ```python theme={null} import requests import json url = "https://api.pdf.co/v1/email/send" payload = json.dumps({ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "from": "John Doe ", "to": "Partner ", "subject": "Check attached sample pdf", "bodytext": "Please check the attached pdf", "bodyHtml": "Please check the attached pdf", "smtpserver": "smtp.gmail.com", "smtpport": "587", "smtpusername": "my@gmail.com", "smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833", "async": False }) headers = { 'Content-Type': 'application/json', 'x-api-key': 'ADD_YOUR_API_KEY' } response = requests.request("POST", url, headers=headers, data=payload) print(response.text) ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.Net; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source PDF file. const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf"; // Email Details const string From = "John Doe "; const string To = "Partner "; const string Subject = "Check attached sample pdf"; const string BodyText = "Please check the attached pdf"; const string BodyHtml = "Please check the attached pdf"; const string SmtpServer = "smtp.gmail.com"; const string SmtpPort = "587"; const string SmtpUserName = "my@gmail.com"; const string SmtpPassword = "app specific password created as https://support.google.com/accounts/answer/185833"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Email Send` API call string url = "https://api.pdf.co/v1/email/send"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileUrl); parameters.Add("from", From); parameters.Add("to", To); parameters.Add("subject", Subject); parameters.Add("bodytext", BodyText); parameters.Add("bodyHtml", BodyHtml); parameters.Add("smtpserver", SmtpServer); parameters.Add("smtpport", SmtpPort); parameters.Add("smtpusername", SmtpUserName); parameters.Add("smtppassword", SmtpPassword); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { Console.WriteLine("Email Sent Successfully!"); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("application/json"); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM) .addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml") .build(); Request request = new Request.Builder() .url("https://api.pdf.co/v1/email/extract-attachments") .method("POST", body) .addHeader("Content-Type", "application/json") .addHeader("x-api-key", "{{x-api-key}}") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} 'https://api.pdf.co/v1/email/send', CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => '', CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => 'POST', CURLOPT_POSTFIELDS =>'{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "from": "John Doe ", "to": "Partner ", "subject": "Check attached sample pdf", "bodytext": "Please check the attached pdf", "bodyHtml": "Please check the attached pdf", "smtpserver": "smtp.gmail.com", "smtpport": "587", "smtpusername": "my@gmail.com", "smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833", "async": false }', CURLOPT_HTTPHEADER => array( 'Content-Type: application/json', 'x-api-key: ADD_YOUR_PDFco_KEY_HERE' ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; ?> ```
# Send Email with File Source: https://developer.pdf.co/api/email/send Send an email. An email can be with or without attachment. **Try it live:** [Send Email with File → API Tester](/api-tester/email/send) — send a real request from your browser. ## `POST /v1/email/send` For converting email to PDF please see [PDF from Email](/api/pdf-from-email). ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `from` | string | *Yes* | - | The "From" field with sender name and email | | `to` | string | *Yes* | - | The "To" field with receiver name and email | | `subject` | string | *Yes* | - | The subject for the outgoing email. | | `bodytext` | string | *No* | - | The plain text version of the outgoing email message. | | `bodyhtml` | string | *No* | - | The HTML version of the outgoing email message. | | `smtpserver` | string | *Yes* | - | The SMTP server to use for sending the email. | | `smtpport` | integer | *Yes* | - | The port number of the SMTP server. | | `smtpusername` | string | *Yes* | - | The username for the SMTP server. | | `smtppassword` | string | *Yes* | - | The password for the SMTP server. If you use Gmail then you need to generate an [app-specific password](https://support.google.com/accounts/answer/185833) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "from": "John Doe ", "to": "Partner ", "subject": "Check attached sample pdf", "bodytext": "Please check the attached pdf", "bodyHtml": "Please check the attached pdf", "smtpserver": "smtp.gmail.com", "smtpport": "587", "smtpusername": "my@gmail.com", "smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "error": false, "status": 200, "remainingCredits": 60095 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/email/send' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "from": "John Doe ", "to": "Partner ", "subject": "Check attached sample pdf", "bodytext": "Please check the attached pdf", "bodyHtml": "Please check the attached pdf", "smtpserver": "smtp.gmail.com", "smtpport": "587", "smtpusername": "my@gmail.com", "smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833", "async": false }' ``` # File Download Source: https://developer.pdf.co/api/file-download Download files from PDF.co Files Storage using a unique filetoken; access is restricted to the account that owns the file. ## `GET /v1/file/download/{filetoken}` **Endpoint URL Format:** ``` https://api.pdf.co/v1/file/download/{filetoken} ``` Replace `{filetoken}` with the actual filetoken identifier. For example: ``` https://api.pdf.co/v1/file/download/a1d30e75adf5eaa................. ``` **Key Features:** * **Exclusive Access:** The endpoint strictly controls access, allowing only the account owner with the correct filetoken to retrieve the associated file. * **Secure File Retrieval:** Files are securely stored and can only be accessed through the authenticated API endpoint using the filetoken. * **File Token Based:** Uses a unique filetoken identifier to access files stored in PDF.co's built-in file storage. Files must be uploaded to PDF.co's built-in file storage at [https://app.pdf.co/files](https://app.pdf.co/files) to obtain a filetoken. The filetoken is used to securely reference and retrieve files through the API. ## Request Headers | Header | Type | Required | Description | | ----------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------ | | `x-api-key` | string | *Yes* | Your API key for authentication. Get your API key by registering at [https://app.pdf.co](https://app.pdf.co) | ## Response The endpoint returns the file content directly with the appropriate Content-Type header based on the file type. The response is a binary file stream. ### Response Headers | Header | Type | Description | | --------------------- | ------- | ---------------------------------------------------------------- | | `Content-Type` | string | The MIME type of the file (e.g., `application/pdf`, `image/png`) | | `Content-Disposition` | string | The filename and disposition information | | `Content-Length` | integer | The size of the file in bytes | ### Error Responses If an error occurs, the endpoint returns a JSON response with the following structure: | Parameter | Type | Description | | ----------- | ------- | ----------------------------------------------------------------------------------------------------------------------- | | `error` | boolean | Indicates whether an error occurred (`true` means error) | | `status` | integer | Status code of the request (200, 404, 403, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `message` | string | Error message describing what went wrong | | `errorCode` | integer | Error code of the request (400, 401, 403, 404, 500, etc.) | ## `Example` Response (Success) On successful request, the endpoint returns the file binary content directly. The following example shows the response headers for a PDF file: **Response Headers:** ``` Content-Type: application/pdf Content-Disposition: attachment; filename="document.pdf" Content-Length: 245678 ``` ## `Example` Error Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "status": "error", "errorCode": 404, "error": true, "message": "record not found. IMPORTANT: If you need to set JSON data then convert it into string first (e.g. using JSON.stringify(obj) ). Check https://developer.pdf.co for more details." } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/file/download/YOUR_FILETOKEN_HERE' \ --header 'x-api-key: *******************' \ --output downloaded-file.pdf ``` ```javascript theme={null} var https = require("https"); var fs = require("fs"); const API_KEY = "*************************************"; const FILETOKEN = "YOUR_FILETOKEN_HERE"; function downloadFile(apiKey, filetoken) { return new Promise((resolve, reject) => { // Prepare request to `file/download` API endpoint let reqOptions = { host: "api.pdf.co", path: `/v1/file/download/${filetoken}`, headers: { "x-api-key": apiKey } }; // Send request https.get(reqOptions, (response) => { if (response.statusCode === 200) { // Create write stream for downloaded file const fileStream = fs.createWriteStream("downloaded-file.pdf"); response.pipe(fileStream); fileStream.on("finish", () => { fileStream.close(); console.log("File downloaded successfully"); resolve(); }); } else { // Handle error response let data = ""; response.on("data", (chunk) => { data += chunk; }); response.on("end", () => { try { const errorData = JSON.parse(data); console.log("Error: " + errorData.message); reject(errorData); } catch (e) { console.log("Error: " + response.statusCode); reject(new Error(`HTTP ${response.statusCode}`)); } }); } }) .on("error", (e) => { // Request error console.log("error: " + e); reject(e); }); }); } downloadFile(API_KEY, FILETOKEN); ``` ```python theme={null} import requests import os # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # File token from PDF.co file storage FILETOKEN = "YOUR_FILETOKEN_HERE" # Destination file path DESTINATION_FILE = "downloaded-file.pdf" # Prepare URL for file download url = f"{BASE_URL}/file/download/{FILETOKEN}" # Execute request response = requests.get(url, headers={"x-api-key": API_KEY}, stream=True) if response.status_code == 200: # Save file to disk with open(DESTINATION_FILE, "wb") as file: for chunk in response.iter_content(chunk_size=8192): file.write(chunk) print(f"File downloaded successfully as '{DESTINATION_FILE}'") else: # Handle error response try: error_data = response.json() print(f"Error: {error_data.get('message', 'Unknown error')}") except: print(f"Error: HTTP {response.status_code}") ``` ```php theme={null} ``` ```csharp theme={null} using System; using System.Net; using System.IO; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; const string FILETOKEN = "YOUR_FILETOKEN_HERE"; const string DESTINATION_FILE = "downloaded-file.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare URL for file download string url = $"https://api.pdf.co/v1/file/download/{FILETOKEN}"; try { // Download file webClient.DownloadFile(url, DESTINATION_FILE); Console.WriteLine($"File downloaded successfully as '{DESTINATION_FILE}'"); } catch (WebException e) { Console.WriteLine($"Error: {e.Message}"); } finally { webClient.Dispose(); } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` # Delete Temporary File Source: https://developer.pdf.co/api/file-upload/delete Deletes temporary file (that was uploaded by you or generated by API). **Try it live:** [Delete Temporary File → API Tester](/api-tester/file-upload/delete) — send a real request from your browser. ## `POST /v1/file/delete` All temporary files are auto removed after 1 hour. You may use [File Upload](/api/file-upload/overview#temporary-files-upload) methods to explicitly force remove temp files once you don't need them. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | --------- | ------ | -------- | ------- | -------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL of the previously uploaded temporary file or output file that was generated by the API method. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `message` | string | Message of the request | | `credits` | integer | Number of credits consumed by the request | | `duration` | integer | Time taken for the operation in milliseconds | | `errorCode` | integer | Error code of the request (400, 401, 402, 403, 404, 500, etc.) | | `remainingCredits` | integer | Number of credits remaining in the account | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "error": false, "status": 200, "remainingCredits": 9999986 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/file/delete' --header 'x-api-key: *******************' --data-raw '{ "url": "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); const API_KEY = "*************************************"; function deleteFile(apiKey, fileName) { return new Promise(resolve => { // Prepare request to `file/delete` API endpoint let queryPath = `/v1/file/delete?url=${fileName}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": apiKey } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.status == 200) { console.log("remainingCredits: " + data.remainingCredits); resolve([data.remainingCredits]); } else { // Service reported error console.log("Error"); } }); }) .on("error", (e) => { // Request error console.log("error: " + e); }); }); } let result = deleteFile(API_KEY, "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf"); ``` ```python theme={null} # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" fileName = "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf" url = "{}/file/delete?url={}".format(BASE_URL, fileName) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["status"] == 200: remainingCredits = json["remainingCredits"] ``` ```php theme={null} ``` # Generate Pre-signed URL Source: https://developer.pdf.co/api/file-upload/generate-presigned-url Generate a presigned URL you can PUT a local file to, returning an accessible link for use with other PDF.co API endpoints. **Try it live:** [Generate Pre-signed URL → API Tester](/api-tester/file-upload/generate-presigned-url) — send a real request from your browser. ## `GET /v1/file/upload/get-presigned-url` With this method you can upload files up to 2GB in size. Please note that to process these files you should use async=true mode with data extraction and tools endpoints along with [Job Check](/api/job-check) to check status of background jobs you create. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `presignedUrl` | string | The presigned URL to upload the file | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "presignedUrl": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/A1VGV42YE0NWXMKEB4BUIWNYGKXEWTND/test.pdf?X-Amz-Expires=900&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAIZJDPLX6D7EHVCKA/20220913/us-west-2/s3/aws4_request&X-Amz-Date=20220913T074159Z&X-Amz-SignedHeaders=content-type;host&X-Amz-Signature=53f326afde5bcfb3b2714ee8cb5322795bf10a03feb7dab3764e6ca63c017f43", "url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/A1VGV42YE0NWXMKEB4BUIWNYGKXEWTND/test.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEBgaDLZTUxFLOwF9iiGk%2FyKCATiLp%2FRn9nPmt%2Fey9PcilcRMXtLl0TS6IFNOpk%2BKtSF%2B%2BEVcbNFThw4c1KVx21RQxT5zf7csSEESGov1Xd4uDhF0xGoVkXff9saXGVUtgKrYgPKhUfv5KEO7gz3E0t%2FqCPZJn2KGs1yMbUkohzeIrEd0NH8EVvqfxrfCcW0ZANiG2iMoh8eAmQYyKLjRMfg02ZJPTgoFPQmfMyYt0FacTg4RhkP3PeD9mrWLefDXCwcYkkI%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHFHQYL4OV/20220913/us-west-2/s3/aws4_request&X-Amz-Date=20220913T074159Z&X-Amz-SignedHeaders=host&X-Amz-Signature=9b1a90f36635459bb40f09b0fc6fe3eba185ba3cfdb0a8ef1096ac9efa9b6299", "error": false, "status": 200, "name": "test.pdf", "credits": 7, "duration": 0, "remainingCredits": 98191146 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET https://api.pdf.co/v1/file/upload/get-presigned-url?name=test.pdf&encrypt=true --header 'x-api-key: YOUR_API_KEY' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); const API_KEY = "**********************************************"; function getPresignedUrl(apiKey) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?name=test.pdf`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": apiKey } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { console.log("presignedUrl: " + data.presignedUrl); // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } let result = getPresignedUrl(API_KEY); ``` ```python theme={null} # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" fileName = "test.pdf" url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] ``` ```php theme={null} ``` # Get MD5 Hash of File by URL Source: https://developer.pdf.co/api/file-upload/hash Calculate the MD5 hash of a file from its URL, useful for detecting whether a source document has been modified between requests. **Try it live:** [Get MD5 Hash of File by URL → API Tester](/api-tester/file-upload/hash) — send a real request from your browser. ## `POST /v1/file/hash` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | --------- | ------ | -------- | ------- | -------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | -------------------------------------------- | | `hash` | string | Hash of the final PDF file stored in S3. | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "hash": "d942e5becdcb0386598cce15e9e56deb1ca9d893b8578a88eca4a62f02c4000b", "remainingCredits": 98143 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/file/hash' --header 'x-api-key: *******************' --data-raw '{ "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf" }' ``` # File Upload Overview Source: https://developer.pdf.co/api/file-upload/overview You can upload files as temporary files into PDF.co. Temporary files are stored for 1 hour by default and then auto removed. To store files permanently (pdf templates, images you want to reuse) please use [PDF.co Built-In Files Storage](https://app.pdf.co/files) instead. You can also use 3rd party cloud services: * **Dropbox**: you can use `public link` to a file from Dropbox. * **Google Drive**: you can use link to a file that was shared as `anyone with a link`. * **Google Docs/Sheets/Slides**: you can use a link to a document in Google Docs that was shared as `anyone with a link`. * Any publicly accessible URL from any cloud service or web source that provides a direct link to the uploaded file. **IMPORTANT NOTE FOR GOOGLE DRIVE/DOCS** users: free Google Drive/Docs limits the number of requests to their files. If you use a link to file or document from Google Drive or Google Drive then make sure you have no more than 5-10 requests per minute. Otherwise Google Drive returns no file or error page. ## Temporary Files Upload You can upload temporary files up to 2GB in size. Please note that to process these files you should use `async=true` mode with data extraction and tools endpoints along with [/job/check](/api/job-check) to check status of background jobs you create. ## Steps to Upload File 1. First, call [/file/upload/get-presigned-url](/api/file-upload/generate-presigned-url). It will generate link for uploading (`presignedUrl`) and final link (`url`). 2. Now send your file to the `presignedUrl` link using the `PUT` method within the next 30 minutes. 3. Once finished, use `url` to access the file you have just uploaded. Note: all uploaded files are considered to be temporary files and are automatically permanently removed after 1 hour. # Upload Small File Source: https://developer.pdf.co/api/file-upload/upload Uploads a small (up to 100KB) local file as a temporary file in PDF.co storage. Note: temporary files are automatically permanently removed after 1 hour. **Try it live:** [Upload Small File → API Tester](/api-tester/file-upload/upload) — send a real request from your browser. ## `POST /v1/file/upload` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/1a4a92ac805c41c28ef75a24e0f35ba5/sample.pdf", "error": false, "status": 200, "name": "sample.pdf", "remainingCredits": 98145 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/file/upload' --header 'x-api-key: *******************' --form 'file=@"/path/to/file"' ``` ```javascript theme={null} var request = require('request'); var fs = require('fs'); var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/file/upload', 'headers': { 'x-api-key': '{{x-api-key}}' }, formData: { 'file': { 'value': fs.createReadStream('/path/to/file'), 'options': { 'filename': 'filename' 'contentType': null } } } }; request(options, function (error, response) { if (error) throw new Error(error); let data = JSON.parse(response.body); console.log(data); }); ``` ```python theme={null} import requests url = "https://api.pdf.co/v1/file/upload" payload = {} files = [ ('file', open('/path/to/file','rb')) ] headers = { 'x-api-key': '{{x-api-key}}' } response = requests.request("POST", url, headers=headers, json = payload, files = files) print(response.text.encode('utf8')) ``` ```php theme={null} "https://api.pdf.co/v1/file/upload", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "POST", CURLOPT_POSTFIELDS => array('file'=> new CURLFILE('/path/to/file')), CURLOPT_HTTPHEADER => array( "x-api-key: {{x-api-key}}" ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; ?> ```
```csharp theme={null} using System; using RestSharp; namespace HelloWorldApplication { class HelloWorld { static void Main(string[] args) { var client = new RestClient("https://api.pdf.co/v1/file/upload"); client.Timeout = -1; var request = new RestRequest(Method.POST); request.AddHeader("x-api-key", "{{x-api-key}}"); request.AddFile("file", "/path/to/file"); IRestResponse response = client.Execute(request); Console.WriteLine(response.Content); } } } ```
# Upload File Using Base64 Source: https://developer.pdf.co/api/file-upload/upload-base64 Upload a file as base64-encoded source data and receive a temporary URL; the file is automatically removed after one hour. **Try it live:** [Upload File Using Base64 → API Tester](/api-tester/file-upload/upload-base64) — send a real request from your browser. ## `POST /v1/file/upload/base64` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------ | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file` | string | *Yes* | - | Base64-encoded file bytes. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/7588d614c9ad41eb98ec317a02abda63/uploadfile.txt", "error": false, "status": 200, "remainingCredits": 77769 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/file/upload/base64' --header 'x-api-key: *******************' --form 'file="data:image/gif;base64,R0lGODlhEAAQAPUtACIiIScnJigoJywsLDIyMjMzMzU1NTc3Nzg4ODk5OTs7Ozw8PEJCQlBQUFRUVFVVVVhYWG1tbXt7fInDRYvESYzFSo/HT5LJVJPJVJTKV5XKWJbKWZbLWpfLW5jLXJrMYaLRbaTScKXScKXScafTdIGBgYODg6alprLYhbvekr3elr3el9Dotf///wAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACH5BAAAAAAAIf8LSW1hZ2VNYWdpY2sNZ2FtbWE9MC40NTQ1NQAsAAAAABAAEAAABpJAFGgkKhpFIRHpw2qBLJiLdCrNTFKt0wjD2Xi/G09l1ZIwRJeNZs3uUFQtEwCCVrM1bnhJYHDU73ktJQELBH5pbW+CAQoIhn94ioMKB46HaoGTB5WPaZmMm5wOIRcekqChliIZFXqoqYYkE2SaoZuWH1gmAgsIvr8ICQUPTRIABgTJyskFAw1ZDBAO09TUDw0RQQA7"' ``` # Upload File via Pre-signed URL Source: https://developer.pdf.co/api/file-upload/upload-presigned-url-put Upload files up to 100MB directly via HTTP PUT to a presigned URL, for use with async-mode data extraction and processing endpoints. `PUT {presigned url}` **Important** The presigned URL must be retreived from the /file/upload/get-presigned-url for the PUT operation to succeed. Content-Type header When sending PUT request don't forget to add Content-Type header with proper value based on input file type. For example: | File Extension | Content-Type Value | | ---------------------- | -------------------------- | | `.txt .csv .xml .json` | text/plain | | `.pdf` | application/pdf | | `.msg .eml` | application/vnd.ms-outlook | | `.doc` | application/msword | If you're not sure then use application/octet-stream header. It works for most file types. All uploaded files are treated as temporary files and are automatically permanently removed after 1 hour. If you have a file that you want to reuse over and over, please upload it to PDF.co Built-In Files Storage and get its filetoken:// link that you may reuse inside PDF.co API. ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "presignedUrl": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c72bf56341142ba83c8f98b47f14d62/test.pdf?X-Amz-Expires=900&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAIZJDPLX6D7EHVCKA/20200302/us-west-2/s3/aws4_request&X-Amz-Date=20200302T143951Z&X-Amz-SignedHeaders=host&X-Amz-Signature=8650913644b6425ba8d52b78634698e5fc8970157d971a96f0279a64f4ba87fc", "url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c72bf56341142ba83c8f98b47f14d62/test.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEGgaDA9KaTOXRjkCdCqSTCKBAW9tReCLk1fVTZBH9exl9VIbP8Gfp1pE9hg6et94IBpNamOaBJ6%2B9Vsa5zxfiddlgA%2BxQ4tpd9gprFAxMzjN7UtjU%2B2gf%2FKbUKc2lfV18D2wXKd1FEhC6kkGJVL5UaoFONG%2Fw2jXfLxe3nCfquMEDo12XzcqIQtNFWXjKPWBkQEvmii4tfTyBTIot4Na%2BAUqkLshH0R7HVKlEBV8btqa0ctBjwzwpWkoU%2BF%2BCtnm8Lm4Eg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHEGHTOA4W/20200302/us-west-2/s3/aws4_request&X-Amz-Date=20200302T143951Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=243419ac4a9a315eebc2db72df0817de6a261a684482bbc897f0e7bb5d202bb9", "error": false, "status": 200, "name": "test.pdf", "remainingCredits": 98145 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request PUT '' --header 'x-api-key: YOUR_API_KEY' --header 'Content-Type: application/octet-stream' --data-binary '@./sample.pdf' ``` ```javascript theme={null} function uploadFile(apiKey, localFile, uploadFileUrl) { return new Promise(resolve => { fs.readFile(localFile, (err, data) => { request({ method: "PUT", url: uploadFileUrl, body: data, headers: { "Content-Type": "application/octet-stream", "x-api-key": apiKey } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + err); } }); }); }); } ``` ```python theme={null} uploadFileUrl = "file URL retrieved from /file/upload/get-presigned-url" with open(fileName, 'rb') as file: requests.put(uploadFileUrl, data=file, headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"}) ``` ```php theme={null} ``` # Upload File from URL [GET] Source: https://developer.pdf.co/api/file-upload/upload-url-get Download a file from a source URL using a GET request and store it as a temporary PDF.co file, auto-deleted after one hour. **Try it live:** [Upload File from URL → API Tester](/api-tester/file-upload/upload-url-get) — send a real request from your browser. ## `GET /v1/file/upload/url` This method do is same as /v1/file/upload/url but using get method. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/97415d1c45a04b29ac42c8dc01883316/sample.pdf", "error": false, "status": 200, "name": "sample.pdf", "remainingCredits": 77765 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/file/upload/url?url=pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf' --header 'x-api-key: ******************' ``` # Upload File from URL [POST] Source: https://developer.pdf.co/api/file-upload/upload-url-post Download a file from a source URL using a POST request and store it as a temporary PDF.co file, auto-deleted after one hour. **Try it live:** [Upload File from URL → API Tester](/api-tester/file-upload/upload-url-post) — send a real request from your browser. ## `POST /v1/file/upload/url` This method do is same as /v1/file/upload/url but using post method. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/1a4a92ac805c41c28ef75a24e0f35ba5/sample.pdf", "error": false, "status": 200, "name": "sample.pdf", "remainingCredits": 98145 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/file/upload' --header 'x-api-key: *******************' --form 'file=@"/path/to/file"' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); const API_KEY = "*************************************"; function upload(apiKey, fileName) { return new Promise(resolve => { // Prepare request to `file/upload/url` API endpoint let queryPath = `/v1/file/upload/url?url=${fileName}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": apiKey } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.status == 200) { console.log("temp url: " + data.url); console.log("remainingCredits: " + data.remainingCredits); resolve([data.remainingCredits]); } else { // Service reported error console.log("Error"); } }); }) .on("error", (e) => { // Request error console.log("error: " + e); }); }); } let result = upload(API_KEY, "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"); ``` ```python theme={null} # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" fileName = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf" url = "{}/file/upload/url?url={}".format(BASE_URL, fileName) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["status"] == 200: temp_url = json["url"] remainingCredits = json["remainingCredits"] ``` ```php theme={null} ``` # PDF Forms Info Reader Source: https://developer.pdf.co/api/forms/info-reader Get information about fillable form fields inside a **PDF** file. **Try it live:** [PDF Forms Info Reader → API Tester](/api-tester/forms/info-reader) — send a real request from your browser. ## `POST /v1/pdf/info/fields` For one-time check of PDF file information and find form fields please use PDF [Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). Extracts information about fillable PDF fields (fillable edit boxes, fillable check-boxes, radio buttons, combo-boxes) from input PDF file along with general information about the input PDF document. The purpose of this endpoint is to get information about fillable PDFs for use with PDF.co [PDF Add](/api/pdf-add) method. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------- | ------ | ------------- | | `info` | object | Info details. | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "info": { "PageCount": 3, "Author": "SE:W:CAR:MP", "Title": "2019 Form 1040", "Producer": "macOS Version 10.15.1 (Build 19B88) Quartz PDFContext", "Subject": "U.S. Individual Income Tax Return", "CreationDate": "8/7/2020 11:17:29 AM", "Bookmarks": "", "Keywords": "Fillable", "Creator": "Adobe LiveCycle Designer ES 9.0", "Encrypted": false, "PasswordProtected": false, "PageRectangle": { "Location": { "IsEmpty": true, "X": 0, "Y": 0 }, "Size": "612, 792", "X": 0, "Y": 0, "Width": 612, "Height": 792, "Left": 0, "Top": 0, "Right": 612, "Bottom": 792, "IsEmpty": false }, "ModificationDate": "8/7/2020 11:17:29 AM", "EncryptionAlgorithm": "None", "PermissionPrinting": true, "PermissionModifyDocument": true, "PermissionContentExtraction": true, "PermissionModifyAnnotations": true, "PermissionFillForms": true, "PermissionAccessibility": true, "PermissionAssemble": true, "PermissionHighQualityPrint": true, "FieldsInfo": { "Fields": [ { "PageIndex": 1, "Type": "CheckBox", "FieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[3]", "Value": "False", "Left": 340.39898681640625, "Top": 67.99798583984375, "Width": 8, "Height": 8 }, { "PageIndex": 1, "Type": "CheckBox", "FieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[4]", "Value": "False", "Left": 441.1990051269531, "Top": 67.99798583984375, "Width": 8, "Height": 8 }, { "PageIndex": 1, "Type": "EditBox", "FieldName": "topmostSubform[0].Page1[0].f1_03[0]", "Value": "", "Left": 238.60000610351562, "Top": 111.9990234375, "Width": 228.39999389648438, "Height": 14.0009765625 }, { "PageIndex": 2, "Type": "EditBox", "FieldName": "topmostSubform[0].Page2[0].PaidPreparer[0].Preparer[0].f2_37[0]", "Value": "", "Left": 509.7449951171875, "Top": 474.0010070800781, "Width": 66.2550048828125, "Height": 11.998992919921875 } ] } }, "error": false, "status": 200, "remainingCredits": 59987 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/info/fields' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "async": false }' ``` ```javascript theme={null} var request = require('request'); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/pdf/info/fields', 'headers': { 'x-api-key': '{{x-api-key}}' }, formData: { 'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf' } }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file url. You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdf-temp-files.s3.amazonaws.com/R2FBM39LFX1BFC860O06XU0TL613JTZ9/f1040-form-filled.pdf " Async = "False" # Destination PDF file name DestinationFile = ".\\result.pdf" parameters = {} parameters["async"] = Async parameters["name"] = os.path.basename(DestinationFile) parameters["url"] = SourceFileURL # Prepare URL for 'Info Fields' API request url = "{}/pdf/info/fields".format(BASE_URL) response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() for field in json["info"]["FieldsInfo"]["Fields"]:print(field["FieldName"] + "=>" + field["Value"]) ``` ```csharp theme={null} using System; using RestSharp; namespace HelloWorldApplication { class HelloWorld { static void Main(string[] args) { var client = new RestClient("https://api.pdf.co/v1/pdf/info/fields"); client.Timeout = -1; var request = new RestRequest(Method.POST); request.AddHeader("x-api-key", "{{x-api-key}}"); request.AlwaysMultipartFormData = true; // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ request.AddParameter("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf"); IRestResponse response = client.Execute(request); Console.WriteLine(response.Content); } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("text/plain"); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM) .addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf") .build(); Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/info/fields") .method("POST", body) .addHeader("x-api-key", "{{x-api-key}}") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} todo "https://api.pdf.co/v1/pdf/info/fields", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "POST", CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf'), CURLOPT_HTTPHEADER => array( "x-api-key: {{x-api-key}}" ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; var request = require('request'); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/pdf/info/fields', 'headers': { 'x-api-key': '{{x-api-key}}' }, formData: { 'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf' } }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ```
# Getting Started Source: https://developer.pdf.co/api/index Introducing the general concepts for using the PDF.co API, authentication methods, response codes and sample code. ## API Reference The PDF.co Web API is REST-based, making it intuitive and easy to use. To prioritize your data’s security and privacy, all requests are securely transmitted using HTTPS. Kindly note, unsecured HTTP connections are not supported. All requests contain the following **base URL**: `https://api.pdf.co/v1` ## Authenticating Your API Request To authenticate you need to add a header named `x-api-key` using your API Key as the value. ```javascript theme={null} "x-api-key": "sample@sample.com_123a4b567c890d123e456f789g01" ``` The key provided above is just a sample and won’t work for actual API calls. Don’t forget to replace it with your real API Key, which you can find in your [PDF.co Dashboard](https://app.pdf.co/), when making requests. ## Response codes After making a request you will receive a response from the **PDF.co** API. A code `200` means the request was successfull, a `400` means there was an error. However there could be other codes - see [the complete list of available response codes](/api/response-codes). ## Sample Code Here is some sample code which would convert a **PDF** to a **CSV** file. ```javascript theme={null} var data = { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf", "lang": "eng", "inline": true, "pages": "0-", "async": false, "name": "result.csv" } fetch('https://api.pdf.co/v1/pdf/convert/to/csv', { method: 'POST', headers: { 'Accept': 'application/json', 'Content-Type': 'application/json', 'x-api-key': 'sample@sample.com_123a4b567c890d123e456f789g01' }, body: JSON.stringify(data) }) .then(response => response.json()) .then(response => console.log(JSON.stringify(response))) ``` ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/csv' \ --header 'Content-Type: application/json' \ --header 'x-api-key: sample@sample.com_123a4b567c890d123e456f789g01' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf", "lang": "eng", "inline": true, "pages": "0-", "async": false, "name": "result.csv" }' ``` # Background & Job Check Source: https://developer.pdf.co/api/job-check Checks the [status](#available-status-values) of a background job that was previously created with PDF.co API. Use this API to check the status of your asynchronous API calls. **Try it live:** [Background & Job Check → API Tester](/api-tester/job-check) — send a real request from your browser. ## `POST /v1/job/check` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | --------- | ------- | -------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `jobId` | string | *Yes* | - | ID of background that was started asynchronously. To start a new async background job, you should set async to true for API methods. | | `force` | boolean | *No* | `false` | Set to true to forcibly check the status of the background job. Intended to be used with really long and heavy background jobs only. | ### Available Status Values * `working` - background job is currently in work or does not exist. * `success` - background job was successfully finished. * `failed` - background job failed for some reason (see `message` for more details). * `aborted` - background job was aborted. * `unknown` - unknown background job id. Available only when force is set to `true` for input request. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `message` | string | Message of the request | | `pageCount` | integer | Number of pages in the PDF document. | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `jobId` | string | Identifier for the job request | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `jobDuration` | integer | Time taken to execute the job in milliseconds | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "jobid": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. 1 ```json theme={null} { "status": "working", "remainingCredits": 60227 } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. 2 ```json theme={null} { "status": "success", "message": "Success", "url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300/L8QYIZQ6KZOITCT0PXUNPM6HKYSP5OIO.json?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzECcaDAbrXwAd1IYG3nZR5yKCAdcavWT%2BuwTotGsad9asqRzowPa1M4BoIWU0M9FqXNJP8xBIQX1Cn7XTq4ZfpklsxcpGE4WcapfHdooi2uR1QWw4kuUlMGGU92uy7pS0RhaGCEL00ES%2BIb%2F5039yyAFklqfAgDlHvi47I7Pp01y6Ua25RzrZGh6ACOd7le%2BXArnbQs4o4ezNqgYyKD%2FCX1I5ZOS0tu0ND0I%2FUWTHp6OR8He9a0dgVXfiMU7pNkwQqwVVFcM%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAZTLLKK5/20231114/us-west-2/s3/aws4_request&X-Amz-Date=20231114T134932Z&X-Amz-SignedHeaders=host&X-Amz-Signature=e5553e080a23fb158c0514f99c9f70be0cb74f764933d712ba628110d4079b4c", "jobId": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300", "credits": 2, "remainingCredits": 1480582, "duration": 33 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/job/check' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "jobid": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; const jobId = "{your_job_id}"; let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "success") { console.log(`Job success!`); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" jobId = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { // Prepare POST request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("jobid", new JsonPrimitive(jobId)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request to `Job Check` API Request request = new Request.Builder() .url("https://api.pdf.co/v1/job/check") .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); } } ``` ```php theme={null} Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Language Support Source: https://developer.pdf.co/api/language-support Reference list of OCR language codes supported by PDF.co for text extraction and searchable PDF generation. ## Supported Language Codes | Code | Description | | ---------- | ------------------------------ | | `afr` | Afrikaans | | `amh` | Amharic | | `ara` | Arabic | | `asm` | Assamese | | `aze` | Azerbaijani | | `aze_cyrl` | Azerbaijani - Cyrillic | | `bel` | Belarusian | | `ben` | Bengali | | `bod` | Tibetan | | `bos` | Bosnian | | `bul` | Bulgarian | | `cat` | Catalan; Valencian | | `ceb` | Cebuano | | `ces` | Czech | | `chi_sim` | Chinese - Simplified | | `chi_tra` | Chinese - Traditional | | `chr` | Cherokee | | `cym` | Welsh | | `dan` | Danish | | `deu` | German | | `dzo` | Dzongkha | | `ell` | Greek, Modern (1453-) | | `eng` | English | | `enm` | English, Middle (1100–1500) | | `epo` | Esperanto | | `est` | Estonian | | `eus` | Basque | | `fas` | Persian | | `fin` | Finnish | | `fra` | French | | `frk` | Frankish | | `frm` | French, Middle (ca. 1400–1600) | | `gle` | Irish | | `glg` | Galician | | `grc` | Greek, Ancient (-1453) | | `guj` | Gujarati | | `hat` | Haitian; Haitian Creole | | `heb` | Hebrew | | `hin` | Hindi | | `hrv` | Croatian | | `hun` | Hungarian | | `iku` | Inuktitut | | `ind` | Indonesian | | `isl` | Icelandic | | `ita` | Italian | | `ita_old` | Italian - Old | | `jav` | Javanese | | `jpn` | Japanese | | `kan` | Kannada | | `kat` | Georgian | | `kat_old` | Georgian - Old | | `kaz` | Kazakh | | `khm` | Central Khmer | | `kir` | Kirghiz; Kyrgyz | | `kor` | Korean | | `kur` | Kurdish | | `lao` | Lao | | `lat` | Latin | | `lav` | Latvian | | `lit` | Lithuanian | | `mal` | Malayalam | | `mar` | Marathi | | `mkd` | Macedonian | | `mlt` | Maltese | | `msa` | Malay | | `mya` | Burmese | | `nep` | Nepali | | `nld` | Dutch; Flemish | | `nor` | Norwegian | | `ori` | Oriya | | `pan` | Panjabi; Punjabi | | `pol` | Polish | | `por` | Portuguese | | `pus` | Pushto; Pashto | | `ron` | Romanian; Moldavian; Moldovan | | `rus` | Russian | | `san` | Sanskrit | | `sin` | Sinhala; Sinhalese | | `slk` | Slovak | | `slv` | Slovenian | | `spa` | Spanish; Castilian | | `spa_old` | Spanish; Castilian - Old | | `sqi` | Albanian | | `srp` | Serbian | | `srp_latn` | Serbian - Latin | | `swa` | Swahili | | `swe` | Swedish | | `syr` | Syriac | | `tam` | Tamil | | `tel` | Telugu | | `tgk` | Tajik | | `tgl` | Tagalog | | `tha` | Thai | | `tir` | Tigrinya | | `tur` | Turkish | | `uig` | Uighur; Uyghur | | `ukr` | Ukrainian | | `urd` | Urdu | | `uzb` | Uzbek | | `uzb_cyrl` | Uzbek - Cyrillic | | `vie` | Vietnamese | | `yid` | Yiddish | # Merge PDF Source: https://developer.pdf.co/api/merge/pdf Merge multiple PDF files into a single PDF document. **Try it live:** [Merge PDF → API Tester](/api-tester/merge/pdf) — send a real request from your browser. ## `POST /v1/pdf/merge` The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------------- | ------- | -------- | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `RenameMatchingFieldsDuringMerge` | boolean | *No* | `true` | This feature enables the renaming of field names during the merging of PDF files which contain forms. If set to false, it will retain the original field names. This is helpful for merged PDF forms with identical field names when the customer wants to auto-fill the identical field names in other pages. | |     `GenerateBookmarks` | boolean | *No* | `false` | This adds bookmarks to the merged document with names assigned to every merged document in the same order: | |     `zipIncludeFilter` | string | *No* | - | You can control which files to include and exclude from input zip files with a profiles. | |     `zipExcludeFilter` | string | *No* | - | zipIncludeFilter and zipExcludeFilter support \* and ? wildcards. | |     `MergedDocumentTitle` | string | *No* | Title of the first document | Specifies a custom title for the merged document. Overrides the title of the first document during the merge process. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Rename Matching Fields This feature enables the renaming of field names during the merging of **PDF** files which contain forms. If set to `false`, it will retain the original field names. This is helpful for merged **PDF** forms with identical field names when the customer wants to auto-fill the identical field names in other pages. ``` { "profiles": "{ 'RenameMatchingFieldsDuringMerge': false }" } ``` ### Generate Bookmarks This adds bookmarks to the merged document with names assigned to every merged document in the same order: ``` { "profiles": "{'GenerateBookmarks': true, 'BookmarkTitles': [ 'BookmarkName1', 'BookmarkName2', 'BookmarkName3' ] }" } ``` ### Include / Exclude from ZIPS You can control which files to include and exclude from input zip files with a `profiles`. ```json theme={null} // include PDF, XLS and XLSX files { "profiles": "{ 'zipIncludeFilter': '*.pdf,*.xls*' }" } ``` ```json theme={null} // exclude DOC, DOCX, XLS and XLSX files { "profiles": "{ 'zipExcludeFilter': '*.doc*,*.xls*' }" } ``` `zipIncludeFilter` and `zipExcludeFilter` support `*` and `?` wildcards. ### Change Document Title You can chnage the document title during a merge with the following: ```json theme={null} { "profiles": "{ 'MergedDocumentTitle': 'New Title' }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/3ec287356c0b4e02b5231354f94086f2/result.pdf", "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 98465 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/merge' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URLs of PDF files to merge // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFiles = [ "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf" ]; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Merge PDF` API endpoint var queryPath = `/v1/pdf/merge`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFiles.join(",") }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "**********************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF files SourceFile_1 = ".\\sample1.pdf" SourceFile_2 = ".\\sample2.pdf" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): UploadedFileUrl_1 = uploadFile(SourceFile_1) UploadedFileUrl_2 = uploadFile(SourceFile_2) if (UploadedFileUrl_1 != None and UploadedFileUrl_2!= None): uploadedFileUrls = "{},{}".format(UploadedFileUrl_1, UploadedFileUrl_2) mergeFiles(uploadedFileUrls, DestinationFile) def mergeFiles(uploadedFileUrls, destinationFile): """Perform Merge using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrls # Prepare URL for 'Merge PDF' API request url = "{}/pdf/merge".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URLs of PDF files to merge // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ static string[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf" }; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare URL for `Merge PDF` API call string url = "https://api.pdf.co/v1/pdf/merge"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", string.Join(",", SourceFiles)); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URLs of PDF files to merge // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf" }; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Merge PDF` API call String query = "https://api.pdf.co/v1/pdf/merge"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), String.join(",", SourceFiles)); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "PDF Merging" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Merge Various Document Type Source: https://developer.pdf.co/api/merge/various-files Merge PDF from two or more PDF, DOC, XLS, images, even ZIP with documents and images into a new PDF. **Try it live:** [Merge Various Document Type → API Tester](/api-tester/merge/various-files) — send a real request from your browser. ## `POST /v1/pdf/merge2` We do not support images in the HEIC format (Apple’s image format) or the WEBP format. The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed. This endpoint is similar to [/pdf/merge](/api/merge/pdf) but it also supports `zip`, `doc`, `docx`, `xls`, `xlsx`, `rtf`, `txt`, `png`, `jpg` files as source. We recommended to use this endpoint in `async: true` mode because it may need to convert source documents to PDF. This endpoint also consumes more credits because of the internal conversions. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls, https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/images-and-documents.zip", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/3ec287356c0b4e02b5231354f94086f2/result.pdf", "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 98465 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/merge2' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls, https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/images-and-documents.zip", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URLs of files to merge. Supports documents, spreadsheets, images as sources. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFiles = [ "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" ]; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Merge Documents` API endpoint var queryPath = `/v1/pdf/merge2`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFiles.join(","), async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); function checkIfJobIsCompleted(jobId, resultFileUrl) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl); }, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source Document files. Supports documents, spreadsheets, images as sources. SourceFile_1 = ".\\sample1.pdf" SourceFile_2 = ".\\sample.docx" # Destination PDF file name DestinationFile = ".\\result.pdf" # (!) Make asynchronous job Async = True def main(args = None): UploadedFileUrl_1 = uploadFile(SourceFile_1) UploadedFileUrl_2 = uploadFile(SourceFile_2) if (UploadedFileUrl_1 != None and UploadedFileUrl_2!= None): uploadedFileUrls = "{},{}".format(UploadedFileUrl_1, UploadedFileUrl_2) mergeFiles(uploadedFileUrls, DestinationFile) def mergeFiles(uploadedFileUrls, destinationFile): """Perform Merge using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {} parameters["async"] = Async parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrls # Prepare URL for 'Merge Document' API request url = "{}/pdf/merge2".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # URL of the result file resultFileUrl = json["url"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.IO; using System.Net; using Newtonsoft.Json.Linq; using System.Threading; using System.Collections.Generic; using Newtonsoft.Json; // Cloud API asynchronous "Merge Document" job example. Supports documents, spreadsheets, images as sources. // Allows to avoid timeout errors when processing huge or scanned PDF documents. namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URLs of PDF files to merge. Supports documents, spreadsheets, images as sources. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ static string[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" }; // Destination PDF file name const string DestinationFile = @".\result.pdf"; // (!) Make asynchronous job const bool Async = true; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare URL for `Merge PDF` API call string url = "https://api.pdf.co/v1/pdf/merge2"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", string.Join(",", SourceFiles)); parameters.Add("async", Async); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Asynchronous job ID string jobId = json["jobId"].ToString(); // URL of generated PDF file that will available after the job completion string resultFileUrl = json["url"].ToString(); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status); if (status == "success") { // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); break; } else if (status == "working") { // Pause for a few seconds Thread.Sleep(3000); } else { Console.WriteLine(status); break; } } while (true); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } static string CheckJobStatus(string jobId) { using (WebClient webClient = new WebClient()) { // Set API Key webClient.Headers.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; string response = webClient.DownloadString(url); JObject json = JObject.Parse(response); return Convert.ToString(json["status"]); } } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URLs of input files to merge. Supports documents, spreadsheets, images as sources. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx" }; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Merge Document` API call String query = "https://api.pdf.co/v1/pdf/merge2"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), String.join(",", SourceFiles)); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "PDF Merging 2" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF Add Source: https://developer.pdf.co/api/pdf-add Add text, images, forms, other PDFs, fill forms, links to external sites and external PDF files. You can update or modify PDF and scanned PDF files. **Try it live:** [PDF Add → API Tester](/api-tester/pdf-add) — send a real request from your browser. ## `POST /v1/pdf/edit/add` Create new PDF forms with fillable edit boxes, checkboxes and other fillable fields. Quickly create configs for [PDF.co](https://pdf.co/) API, [Zapier](https://zapier.com/), [Make with PDF.co PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). To add a signature, save the signature as an image and insert it using the images attribute. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `password` | string | *No* | - | Password for the PDF file. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `inline` | boolean | *No* | `false` | Set to `true` to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `annotations` | array\[object] | *No* | - | | |      `text` | string | *Yes* | - | String to add, if you need to insert a line break then use \n or `{{$$newLine}}`. You can also use built-in macros like `{{$$PageNumber}}` and custom data macros. | |      `x` | integer | *Yes* | - | X coordinate (zero point is in the top left corner). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |      `y` | integer | *Yes* | - | Y coordinate (zero point is in the top left corner). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |      `type` | string | *No* | `Text` | Set object type. Available types: `text` = text object (default), `textField` = text input field, `TextFieldMultiline` = multiline fillable text field, `checkbox` = checkbox field. | |      `id` | string | *No* | - | Sets id of the form field if type is not text. | |      `leading` | integer | *No* | - | Sets a custom line height for text. The value defines the vertical spacing between lines. Larger values increase the space between lines, while smaller values tighten the spacing. | |      `width` | integer | *No* | - | Width of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. | |      `height` | integer | *No* | - | Height of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. | |      `alignment` | string | *No* | `left` | Sets text alignment within the width of the text box. Valid values: left, center, right. Default is left. | |      `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. | |      `color` | string | *No* | `#000000` | Sets the text color. Default is `#000000` (non-transparent black). Color in **RRGGBB** or **AARRGGBB** format where **AA** is the transparency component. For example, 50% transparent green is `#8000FF00`. | |      `link` | string | *No* | - | Sets link on click for text. | |      `size` | integer | *No* | `12` | Sets font size. | |      `transparent` | boolean | *No* | `true` | Set to `false` to force disable any transparency and draw a white background under the text. | |      `fontName` | string | *No* | `Arial` | Set font name to use. Default is "Arial". See [availabe fonts](#available-fonts). | |      `fontBold` | boolean | *No* | `false` | Set to `true` to enable bold font style. | |      `fontStrikeout` | boolean | *No* | `false` | Set to `true` to enable strikeout font style. | |      `fontUnderline` | boolean | *No* | `false` | Set to `true` to enable underline font style. | |      `RotationAngle` | integer | *No* | `0` | Set rotation angle in degrees. Default is `0` degrees. | | `images` | array\[object] | *No* | - | | |      `url` | string | *Yes* | - | URL to image or PDF as HTTP link, file token, or datauri:.. URL (with base64 encoded image). | |      `x` | integer | *Yes* | - | X coordinate (zero point is in the top left corner). Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |      `y` | integer | *Yes* | - | Y coordinate (zero point is in the top left corner). Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. | |      `width` | integer | *No* | - | Width of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. | |      `height` | integer | *No* | - | Height of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. | |      `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. | |      `link` | string | *No* | - | Sets link on click for text. | |      `keepAspectRatio` | boolean | *No* | `true` | Set to `false` if don’t need to keep the aspect ratio for the image or PDF added. In this case, it will use the width and height parameters provided. | | `fields` | array\[object] | *No* | - | | |      `fieldName` | string | *Yes* | - | Name of the form field. To find form fields please use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). | |      `text` | string | *Yes* | - | Value to set for this field. If you have a checkbox, set X, true, 1, or another text which is different from false to enable the checkbox. For radio buttons and combo boxes, you need to set the item value in text or index of the item to select. To find form fields please use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). | |      `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. | |      `size` | integer | *No* | - | Override the font size of the text inside the given input field. | |      `fontName` | string | *No* | - | Name of the font to use to fill out the input field. | |      `fontBold` | boolean | *No* | - | Override font bold style of the text input field. | |      `fontItalic` | boolean | *No* | - | Override font italic style of the text input field. | |      `fontStrikeout` | boolean | *No* | - | Override font strikeout style of the text input field. | |      `fontUnderline` | boolean | *No* | - | Override font underline style of the text input field. | | `annotationsString` | string | *No* | - | This parameter represents one or more text objects to add to a PDF. Each object is made of parameter separated by the `;` symbol. | | `imagesString` | string | *No* | - | Adds one or more images or other PDF objects on top of the source PDF. Each object is made of parameter separated by the `;` symbol. | | [`fieldsString`](#fieldsstring) | string | *No* | - | Set values for fillable PDF field objects. Each object is made of parameter separated by the `;` symbol. See [fieldsString](#fieldsstring) for more information. | | `profiles` | object | *No* | - | - | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `Pages[0].SetCropBox()` | array\[string] | *No* | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). | |     `DisableLigatures` | boolean | *No* | `false` | To disable ligaturization, for example for Hebrew, use the following: | |     `FlattenDocument()` | boolean | *No* | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### fieldsString Set values for fillable PDF fields (i.e. fill pdf fields in pdf forms). To fill fields in PDF form, use the following format: `page;fieldName;value`. Also, the advanced format can be used to override font name, size and style: `0;fieldName;Field Text;12+bold+italic+underline+strikeout;FontName` Where: `0;editbox1;text is here;12+bold;Arial` To fill the checkbox, use true, for example: `0;checkbox1;true`. To separate multiple objects, use the `|` separator. To get the list of all fillable fields in a PDF form, use the :ref:post-tag-pdf-info-fields endpoint. If you need to include a pipe character within a value, escape it by prefixing with `\\` (for example, `|` should be written as `\\|` and in a JSON string as `\\|`). ### Crop a PDF File Crop a **PDF** file using an array to define the crop area. The crop box is defined by a rectangle `[x, y, width, height]` in **PDF** points (1 Point = 1/72 inches). An A4 page size in points is **595 x 842** ```json theme={null} { "profiles": "{ 'Pages[0].SetCropBox()': ['28', '28', '539', '786'] }" } ``` ### Disable Ligaturization To disable ligaturization, for example for Hebrew, use the following: ```json theme={null} { "profiles": "{ 'DisableLigatures': true }" } ``` ### Flatten Document Flattening a document renders it as *read-only*. Handy if you want to remove editing or copying capability. ```json theme={null} { "profiles": "{ 'FlattenDocument()': [] }" } ``` ## Available fonts ### Standard Fonts * Arial * Arial Black * Aptos * Aptos Display * Aptos Narrow * Bahnschrift * Calibri * Cambria * Cambria Math * Candara * Comic Sans MS * Consolas * Constantia * Corbel * Courier New * Ebrima * Franklin Gothic Medium * Gabriola * Gadugi * Georgia * HoloLens MDL2 Assets * Impact * Ink Free * Javanese Text * Leelawadee UI * Lucida Console * Lucida Sans Unicode * Malgun Gothic * Marlett * Microsoft Himalaya * Microsoft JhengHei * Microsoft New Tai Lue * Microsoft PhagsPa * Microsoft Sans Serif * Microsoft Tai Le * Microsoft YaHei * Montserrat * Microsoft Yi Baiti * MingLiU-ExtB * Mongolian Baiti * MS Gothic * MV Boli * Myanmar Text * OCR A * OCR A Extended * OCR B * OCR B E * OCR B F * OCR B L * OCR B S * OCR B X * Nirmala UI * Palatino Linotype * Segoe MDL2 Assets * Segoe Print * Segoe Script * Segoe UI * Segoe UI Historic * Segoe UI Emoji * Segoe UI Symbol * SimSun * Sitka Banner * Sitka Banner Semibold * Sitka Display * Sitka Display Semibold * Sitka Heading * Sitka Heading Semibold * Sitka Small * Sitka Small Semibold * Sitka Subheading * Sitka Subheading Semibold * Sitka Text * Sitka Text Semibold * Sylfaen * Symbol * Tahoma * Times New Roman * Trebuchet MS * Verdana * Webdings * Wingdings * Yu Gothic ### Japanese Fonts * MS Gothic * MS Mincho * Yu Gothic ### Chinese Fonts * SimSun * MingLiU * Microsoft YaHei ### Korean Fonts * Malgun Gothic ### Hebrew Fonts * Miriam ### Arabic Fonts * Aldhabi * Andalus * Arabic Typesetting ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `hash` | string | Hash of the final PDF file stored in S3. | | `url` | string | Direct URL to the final PDF file stored in S3. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `pageCount` | integer | Number of pages in the PDF document. | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `remainingCredits` | integer | Number of credits remaining in the account | | `credits` | integer | Number of credits consumed by the request | | `duration` | integer | Time taken for the operation in milliseconds | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | ## `Example` Payload (A) To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "async": false, "inline": true, "name": "f1040-form-filled", "url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true", "imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200", "fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789" } ``` ## `Example` Response (A) To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c336bfcef1a473d98492bda25d8da03/newDocument.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEO7%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDHWK1dY4d4lOgsheliKBATwE%2FZewASPTEnPxTn%2BOdYhP4h3gljAJfqbRvQptDX7wdWLmrBS7Tg4qTU6pAbxIdXChGPjBWpSbtiADJKmqkmyhkUmE8GSM1%2FGtJO6bga2pgzvFLXmzxjTf3%2BFNqwYOvbyApIZdVLoPpEKY6PlCflQtLTd30dhelm6xpB8pitbdhSjdz8KCBjIobVy%2Fjwybwp6OQgB%2FT6QkIo2dU07gtFREdn5jhRyvnS5lkccweBV1%2Bw%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHMV5P3JOS/20210316/us-west-2/s3/aws4_request&X-Amz-Date=20210316T124309Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=95287bf3c007fed4c2c5aeea1ce75c846cc6c68b22aaf35175ebe41a105f54e1", "pageCount": 1, "error": false, "status": 200, "name": "newDocument", "remainingCredits": 9913694, "credits": 3 } ``` ## `Example` Payload (B) To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "async": false, "inline": true, "name": "newDocument", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf", "annotations": [ { "text": "Sample Text 1", "x": 150, "y": 100, "size": 20, "pages": "0-" }, { "text": "sample text that is centered (can also set right or left alignment) ", "x": "10", "y": "10", "width": "500", "height": "200", "size": "7", "pages": "0", "alignment": "center" }, { "text": "Sample Text 2 - Click here to test link\r\n(CLICK ME!)", "x": 250, "y": 240, "size": 24, "pages": "0-", "color": "CCBBAA", "link": "https://pdf.co/", "fontName": "Comic Sans MS", "fontItalic": true, "fontBold": true, "fontStrikeout": false, "fontUnderline": true }, { "text": "Simple text 3", "x": 100, "y": 230, "size": 12, "pages": "0-", "type": "Text" }, { "text": "sample text 3 - input text field", "x": 100, "y": 170, "size": 16, "pages": "0-", "type": "TextField", "id": "textfield1" }, { "x": 200, "y": 120, "size": 16, "pages": "0-", "type": "Checkbox", "id": "checkbox2" }, { "x": 200, "y": 140, "size": 16, "pages": "0-", "type": "CheckboxChecked", "id": "checkbox3" } ], "images": [ { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png", "x": 270, "y": 150, "width": 159, "height": 43, "pages": "0" }, { "url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAgMAAAEtCAYAAACVlWOMAAAgAElEQVR4Xu3dCXxkVZn38f9zK72wiCjdgEInadx1RnFkHDckCQiiIi7grqDMNEmQAXV05p1XBcdlXEFHuxN6FFF5RwUXcGNPAoq7CKKOG3SSRhS6W5ul6Sbpus/7OZWq5FZ1JankJuncvr/6fOYzM6Tuued879N1n3vuWUx8EEAAAQQQQCDXApbr1tN4BBBAAAEEEBDJAEGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEJgDgaNPH3piHOlquX66fMnIG6745GPunYNiKQKBBREgGVgQ5t15Erf2M+54ujx+o1zHSDpU0rJyjR6QdLvcbzDTt7YvLwz84PxV23dnbTn3wgscs+a2hxablrxWsb9SpidKWlGuRSzTJne7ydyvtNi+3bd+1e2SeW0tj1wz+Iimgt1o0taouPMl165/1PDCt2T3nrH9jI3/oDi+StLeJj+xr6f1it1bI86OQOMCJAONW2Xum6UnFdMXZPq7ROUflGuTTC6XybSyKjlwfWrfeMl7vrn+kSFR4NOAQPlG+F3JruzvWXVmvZtlA8Us+FdOPvlXS7es2OcdLvs/4QaWqMBmSZWkcK9EchDSgFststOuW9f8k2SFj+38w4Gj0ZLr5Xq8ZGv7e5rfvOAN2s0nTCQDD5X7Sf29rV/dzVXi9Ag0LEAy0DBVtr7Y3jX0NEnfGfshtxvM4/du36vwg9on/3BD2HzAPq+T2UckPbzUStcNvjw+ceDjq7dmq9W7p7YdnRsPcfPvS/FexULh6Bs+terW3VOTxs/6tDU/XfLQwsp1Lr1J0la5f9yXNH124JOH/LE2mWnr3vho8/jjko6XFEl60Mze2Leu+YvJMz63c/CpBelFkv22v7f50qwkRY2rTf1NkoG5kqSc3SFAMrA71Of5nOGHfr/CgV+S/AQ3WzOwbtXnpvthLnUVR00fkZVuDgW5zurvbfmvea7qHlH8RDLgKyPzF163rrV/soa1dW5oNSu8RuZPlOsRkv/Q3f97oHf14EJiHN01fGQsv0KyW6MoetF1aw/dMvX53do7Nx4r83WSDpPpN0vi0aOu7n303QtZ78V8rmPW3NZcjJpulOnQrPYMLJb4XMzXeU+tG8nAHnhljztt48NHl8bXuuwAc3tWX++qPzbazNKPQRS9Qu6/6u9p/XYjNwgzP9+lx0j6o5l9RbGdP5NzNlq3xfq9RDLQPNlNoOOM21s8LqxNPF0nm1M016cO2HL/Oy699EkjC9HO9s7BN8nsM5LO6e9p+Y9Gz1nqSVqxz/NMeqpFTT3TJxGNlrx7vnfCmjv33lYYeavLzpa0r6RfmmztPsWmL8/0VVl1HGQrmV5s8bl7oiHfZyUZ2AOv/3gy4FpZiHc+e74Gc008XWqfGsYHTP6fB2ze9uGFurntzss4XTLQ0TV4vMu+IOkAme6y2D8dR/6VUGfz6K2SXhP+T7mf29/b8r7penHmoq3jyYDZe/vXNb97LsrMYhkdXUMfdOlf69T9dpdOH+hpubbRdj339I2PKUTx9VLo8Zk8yWrvuv1xRY/2vqG39eeNlj2f31uM8Tmf7aXs+gIkA3tkZLh1dA2vL70PnscbTHvXUBh49k5JF0t+vbme6JH9o1wHlVkvHXlw+2k3Xvj4+/ZI5nKjpkoGOroHn+Vu35D0MLk+uuvgTLf27o1d4Z29XHd5FLUPrFv1h/n2auvecLh5FG50fy1Ix13b03L7XJzz+DN/v9+O0SXnWhT9pG/dqi/NV2JTeqJvGn2tXG936VGSRt10Q6Gos6+7oOXXjbSlrfvufS3e/m2ZtUrxp2R2l1wvkPSS8qDaB9111kBv8/pG2lE1ZmCSZKA0/iKO+xXp/vl4zdK25s4VVti5xuSvc+mxpVd+UhgM/EN3+++VW+77WjJBX6zx2cj14ztzK0AyMLeei6a0xFN700x+0GbSgNBlvGnlyqUD6w68v3Jc+JG+Lxp9u5lCohCmMH5+xeb7/2m+eghK3dYr932VvPR01yrZZnP/ehTvPOfa9Y+6Zybtme13q54IE6PISzfGnUvDIM5nTpWUBbP7C6PflNThbq8e6G3+0mzr0uhxiXElL3PTD5ZEdtI1n2q+s9HjJ/teIjHauLxp5AXTzbV/5ls27rX3A/Fz4shXlgYe9jTfNNWNt+6A1+rK3GPub+zrbf16I20JCcHKTZtGqm6QpVc6Ua9kz59ssGS9sstP2JdLWjJZz0B759A/y/SJ0iDdaK8XJv/tTFXf0JtgUdNBfWtXhVkr41M7n/fm4UeOFvX0QuzbYtO/ytRWTgDqF+e6vLIGwmKOz0auHd+ZWwGSgbn1XESlhQFfQ++U2bmlSrm+6VHcObBu9Z/nv5JuHZ3Dr3PTf5d+GOv2Trg9t3Po8CYV7g7jC47u2vjkovm55n5Eub4bFNmXfWfTJQPrHxmmuu3yaesaOsakC0oD2nb93B3JTrqup/m7893eyZ4I27qGTjHpQjf9SEvjF0w1O6O9e/g/5P6umb7DT9O2Y7qGDitKYV78oyX9Ra6uFVuav3rppVacbbmNJgOlAauFpg9LOiUxtTU26SMHbL7/3ZWb8/jTu/yXbvZjk8IrjXC9HzT5Fe76ZuS2MY4U4uZfyjNitpj5i/vWtX5/tu0YG0sweoFLr5NUp7xSj87Zcg8Dbh8q6SGS9p/kfGEcyJ8U2VkW+0nlMr/c39PyqkbqV1nDQdJ+bvExA+tW31w5rtw794Hxclw3mfztUuG3kcXLdro/MzK9yGXPK72mkkYrayAs9vhsxIbvzJ0AycDcWS7CksZvymFaWJg2GLoLP1Eo7vzQ/D81VyUj9ymKjutfu+pHFaTKU4lJD3f3L8rs/yZuCknLB2T2zv51qz5eeSIqzY8/YJ93u9m/JZ6CtrrsB5Lva9LTy2Wlvik0clHrJQPhiXfZg/Hl5jo6vK4Z6Gn53FRltXcNhZvcexYyGQj1CU+WO4u+fnxgo+smSf/c39v8/Ua6xmvbNJEM6M5CcfT59eLsuW/e+LeFYvErkoVu7NLCRuUEMKx5Ebr7XzewrqU0piLx9PrsxA3vBisU39C39rCh5PnD07NU+FYpuXHd5PGS4yZLJBu5rm1nb9jfdkSXy/RcSTcmezrGelZWXCDZGxspK3zHXF/xSGGdj9fO5DpPNTAxETfhFFP0wrm1dQ8eZG5PK9qSX4wuizdnIT4bteV76QVIBtIbLvoSyk9hoRv9rPLiMg/I/SJT4QPzOeq/9P4yGr2qtOiR6/IVW+5/ReWJr+6PvBSb61cuXSH5/SY72U1PKj3NlOe1P/tNv3nI0qXLL5DZK8tz3m9X7Geu+EvLVZUn2rGbTRy6a1dLNrC86cETp+6udjv69OEnxOYvl+wYWThOYbGdpvBO3eQ/cuniFZu3XVPvdUe9ZGD81YGrON0gzqPPuOOAOI6/Jfkz5H5af2/rhQsbVCFxG36WpP+qLFAVFheS/JwDtrR8YyY9BYlkQPVmsiR6I8LT/RUWFc+o3NQ7Oode76Yww+HHlRtvbZyEVxpT9bJ0dA52uJXGaOzlprcOrGv5RBrL6vJ8zcC61lC/8setY83Gw7zgR8ssrD64l8yOlnRgiDuLdI6Kdlv4chztLK7c1Lpp84rhkPTOKOlLGph0cV9Py+srNWjvHHy5wgwe6a8ht+vvaflZI+3NVnw20iK+k1aAZCCtYIaOH/uhLr5XsjeUn6iLcr/Eo8K7Zzdoze3IM/+8YunIyOHhna+ZBmP57eFHr3IDae8a/JhkYcT8tkh2fKXbvk4ysMtgrZNP9sKWlcPvclfoPv+r5KdItkbSi8vskw5QbD998AWK7GvhNcXkT+aVufPxf5WfUqe+mqWZAHp7X2/zxcmn5nrJQOL98fiNrV7hYZVIN33JTX9r0v/GFncszKucXWtT8l4x/DKXPjjx6sWGZf6uFZvu/1Ij4z4ST7Gh27zq5pR40n6Om61buem+tyXLnJgSq8dWYqUmTkbN9PK+dS1hfEXdT3kxpQtLXfGmn/rS+HlTvZ4JvUx3HbTPIYWdYZVOKxQjvzV2+/P3elrCzVWJOj11uvKSqzDW3rTHb97lHqDJ/l6vUUkDN13z4LLoxMriYRPJgA3PZBpxFuMzQz+1mawqyUAmL1u6SrefMfQkxf5RyY4tP10X5fqWR9G/NJIUhJ6GOGp6j5tCF+l+dWpTNOkXsflauf2NSWEOt0z6dF9P85pwI93liU/20YGeVe+o7ZoeGzC27yUynVhznikHJiYHyEnq27e45ITkvPHSO+FodK2bQmIUVtUrPcBJ2ihpIOzXINMOWXSQPPQY6BnlBCqW/HMrNm/rrNzI6iUDiSe2qu7lShvGRn2Pvj/xznxBXmk0EjljNjtPcfP3ja9KabpLsf/7ii3bLp4qKUgkA4cqspf1r20OPTSlT1v30FnmOi+siFmvt6YmJkrrH9T8tw07i/7s765v/dNU7UgMnl06WfIQxqjEHp83xYC7B9zsKnNd5PJ3mvT3Yz1UkycjVd35Ut0xAZW4qL2pT9WeGoOqeJqIPbtnJslAluOzkRjmOzMXIBmYudkecsSuXcOSQlJwYSHe+fbJxhS0dw4fJ/PPj3WFjn/CgLO7SoPQFAaf+eMnef8/VLSmI29Yd8jGxDv1MLBpS+3AqCRyzc02/GmDFe15feubS12wk306uodOcFdYH/6B5JiFXd/32rC7f0DL4y9P9hRZunk37fyQ3E8NyYNLvQM9zd0heambDHQPvUGuME7gRjNb6/K3yu0WyW8z6WSXnpwY73B7FPvrrrug9QeLKbjqvF4K1btdbt39vauunmrDorFXNPr3/p6W/wwHjQ2C04Bkh5j7i/t6W/vqtbW9ayjMpHilS9+8r7jp5Qcue+he5RkZYcxA3cSqtpxSD8RIdI1cRyQT0PC9ScabhD+F8TRbJL9TsoMlrUokiROncP9if2/La+u1vZFk4OjTB58ZR3Zl6Omq/FuY7prXlFtlMLHqoS+byVLY7XtAfE7nxt9nJkAyMDOvRfvtsIJY7E3PGFjX/OWZVLJ+17A2lrvWqxZcScxJDqOSw4/nxcVC9KmD7jr018n3yuWn+ZfIFG4EyZH+4yOZQx0rP/yS/XCywWbhe+Hm/ZDCyq+adMJY2/y8/p7Wt03XzqrNcxLLKydeIYSpjzNYC6FqUOT2yk2tKhkoL+KTfPKS+4UyCzMrKj0Qoeph5Pxtbnpfo13w07V3qr+H67xpxcYToqJunS6Jqi2nzuulWLLPjjz4wFtq15CoeYodfzqujFwPvS61vTRViV85Gagsd7wl3vrXxLVvKBkYi63y66mqZZOrrl8k2bDcP7Yz9ktrexvKPTfh9VZlnM1YNadYhrmRZCAxM+DQRnc2rCq35vyJ1xh/02h5JZ+JsQY3Lob4TBPbHDs3AiQDc+O4W0upDNRz0x/D09TP1h8xOtMK1ekafsDdTqvMeU8+bUl+ZVMhOm26eeml+fPRyEdk1pm4EY4vf9vRNfSFRqdZJZbPVePrvpdmU1zippMq72hrXh/cUzvLYTq3qtcW5afE0hbRY1vXhilmpRtg4p3sHebRkW7xwZI/rnQ/MQ0uK4z+cro5+NPVZSZ/b+sePM3cPhGZnzDV3glTlbnL6yXXdb48PinZm1KdDEwkeYlrPeXyxxMJov5UjKOjbrhg1e8bTRqrkoozhk9U7GHMyPhMluTaG3J//4ot2z403TiIjjXDj1LBL3XpqeXyJ42ZRpKBmsR2fX9Py+nTXcfy7IvrJAtrMdSMDZiI8UaT5HC+xRaf0xnw9/kXIBmYf+N5P0N7+YfPpFuWjETHXPWZVX+Z7UnLe9ufV+4O/2tkdnzYrnb8x8N160ymbI0tTPSQj5l7dzkhGH9aTEyLmnbOddUP4gw2URo/R3mRl1ijDyv4zrD2QIukGScD5aeq0rr+Jv08eI8s12Nqk4HEQkR7zzThmO21m+y4iZX29Nz0sxXGNywKPR2rJP/svcXNp08koMmb09gNfZ+lO+4a7+qfYmvfmhvl+LUZX4PBdcd0MzMqBuVdO68JKz9WFnJq7xoKa1Ks8UnGp0zmF7ri40LTZeMJwSRtSCYDYRphX2/zK+q9TphIbP13O4tqm3oMhFtb18YPmzysoRBSyV0GCo6vNdDAgMlKGxdTfM51vFPe7ARIBmbntqiOSmT5Ve/GZ1vJqhHU5afftq7hN5h0kWaxln15hb3wlHZccuBUIhmYtvs3uUpfvQGBk94IO4dfZeZhq93SOXbEy5+QuHFXvbZo1Kt2BPdO194Ta9KPvcJI9kCY9KG+npawJsJu+VSNz5jinfdMKpfoqdkloUr0AmwPuzgWl/rPK+/wp+rVqXqtEypTvumOr9onlcprpGej5in9HD+o+QP25+GvyNQexf78mY7PaD/99qcoKlxdGiszSTI61bv9pO34ksSmQ939XwZ6Wz82mX3Nq7nwtV0MEv/+dzbqs5jicyZxx3fnT4BkYP5sF6zksa7ZZZdL3ib3L67Ysu3U6bo/p6tcort27CY6uuR5Y/OZG3tfX1t+4r36LytzyCvJQOUJe7oejUTyMOWAw+S5EzetUjtGRpY8qTyAa2wWhPtJxWjJj5cUdz7+gb2j71WmbE3l09Y5+DYz+2gYyBhGty+JCpGbf1/y5uRiMok56jvSrog33fWa7u9t3UMnmYc9JLRDrhf297bcON0xU/29apxEzZNy1UI4rrOWLxm5qNIzYK5P9PU2v6XeE3PbWOIWBqeG5XzHk4GqZX4b7BWqNze/HNMvbvSGuUsMT7Mw1FSj/qvLqnrav1tx8dj+Cw67pfZ8NStEhlUMwz4Dhdp/g8/t/uOqid6uxv99Lqb4TBOLHDs3AiQDc+O420up3nAk7e53VV29pZvotpHlB5WffpdaZC/rW9t8w8wa7dbeNfRRc3v4PfGmNaFbeeKmUd312dE1dHIs/3dJh5ii29yLVw30rj63ZpfE8ZHqU960Jn7A6/UMhBvOaSpEW8rvl+8312fjeMn76q1cVx4D8TaNrXy4d6WXY6/teni9ZKBm1sLGKPZXzvSJdGbGk397bOvhfUPXfphK+Ye0mxMlRsXvV/u0X76ph96Y8cGeid0B606hLJcXBr+GUfzlQ8e2AU7u/TBV93uy9fUGMlYSQ5e+W4gKL53p9stt3RsOlkeXyG1dvf0jGk8GpNIsgELhmrC+RVhfwlV8aX/PYb+ttCH5aqKyd0Rxpx89tiiTb0i+Xqh6vTKDVwWLKT7nKs4pZ/YCJAOzt1t0R7adPnyERaVBU4eE1d3corMbWTegtiFtazY83gpR2GBn9UQXd9XTTJhK+D9RrA9ed0Hz/062bG34cRzduXRF0T0s2BJG7sut8JeVm+7tCz0XiSfI8dXT2ro2vCRS4RSXh53jkp/1+xaXvKWyoU+jy7nW9nBs8+UthWKxPBirVPw5heLOiyo/zOUTjq97X1prwO0ppXUGrLT+/d6V7yj2l/Vf0Pqdmu7tqgFyNUvahqmbA7LovSs2H/q9mazsNxfBVjN+Y6vc3z7dugH1zluagXLA0Hs8LCHtuqN2p8WqRKE8oLK8VHCYTvjIykyUsNOlLNpf8hfKdVydDXZKlsnXHI32ItVLBqpWxBybBvt+LzZ9fvIli8cW1SoUdxxibk+sWBRd/1t/++Fdk+ipBol2dA+/2t0/W56G+xdz73Gzr5dWM3R/b2mNhzDWZXl8YhikmdxYyE2vrCzZHOqVmK2xy9LfU8XOYorPuYhxypi9AMnA7O0W5ZE1a82HhX5uc3lYsa2vUCzeumRZvLXeD1T4wV3yYLG1ye0Ul7rKiwlVPUFOslBPmGJ4l8lvcdm9Mj1RsQ6WKawzX0oAkp/wYx4XlxwbfoDLXdfhaXD8B6yta/A7Jju+Hu6O5dHee40Un+Bu7/WinTNwQfNPp7oIySemyhPlkWuGDm4qWOgiD/Pgw6c0eDGRSE08mU5e+ANmOrtvXfOnd1lAqc5ywnWWUA4lPyjpNpnCnvZbFOs2RfYUuTeZfD9XKQEJyyEPu8evG+hdPTg3AbfrfhUuXWdmV7qKYWOfPydXkJw459iNcUlx5Bmx610mPa30tzqbUCWmzwXj8cGhHZ2DL3WzcPMLsy7qfcKGVCHZqiRc44lV4tXM+CyDqTzqLWAUvl/a/U+Fr7v0hMTx4bx/lukWuZpM+js37ScvxXByOmg4JCyZ/cm+3pbSQlq1n8SKm9OOgwmpcVvn0GvMLOwNUWnzeJH1dpOcKN++dm/x7ldVBm5Wj0OY2c6Xiys+5ybKKWXmAiQDMzfLwBFuHWcMPdWL9uFptzSdvDU/tqK9pnZOemm++sqNrzf5B+U6qEGMsBDR9+T65Iot2745sSPdhoMjj/pcviIsmBIV4zdVVit0d5mF8Ay7tZbD1HVTf2/L2E2ogU/VDaH8rrlmsaNQ/PhWsmODFEdeYm4vdSvd7MLNOCyrGz73ybXBI//86I4dlyTn1yeSjudPNtd7bJ7/8KtN9v7y2ILpWjDtAlDTFTDV30s3gOV7haTvrTO4jskiH5D83BWbW86r7eGomr7putwPbj5p4FzbGQ4++ow7HhvHxfMkhe2BwzvwUm+JFfwdfWtbft7WOfTW8niM8PWJZKC7EitqbeSd/1RP0ZPsmDgd573m+ppZdP51PYfeOllv2ESCa3c0uiJgWCPE40LoCQgrXYakYIvJP75Pcel5yVUzQwUTY2BqkiK3jq7h9S79Y6O9ZskGL7b4nO5i8Pe5FyAZmHvTRVVi6AbUDnth2NjHpCOn2GY11Hurm/WH+ejTdWOHH4/NB95xhC6CB8UAABcySURBVIrxyxSpo/Sud+JJauypV36lmb66fVnh55MNzAs356U71Gbm/2TuL60kAS6/IpLucVnVNq/9PS0Nx2zipvS4QnHnC65d/6jh0g9qeYpZ+UI1tMTtdBe1PIL+yOlWlSu5HTD8DJP+0c06JA+vdMJNcWz3vrBRU1ikaLl/e6o19aerT6N/D/W566A7nlgoFk8NPTIuhZ0EQ33qfcIy079z2cVebFo/1Y6A5YWdLjXZm/t6mkNvQNUnXPcw1mLZkh33JXuqal4xVM3DH+u9if+fXP/e39saVpac9FPa/MmLV8ptiVvx+fX2eyg/ER8rsxNMepar1BNQ2YZ4q1y/kPTVqFC48uF3H3JbI691JlZajJY3mgw0eq1KsVtaSlxhMbCwF8hLk/s0tHVvONw8+pIrPnWgZ/UPZ1Ju5buLLT5n0waOmZ1Awz+ssyueoxabQHhv/OeH7XNAGAGfrFvtj/JC1bu9c/A8mb2lcr7SZj0edw/0rh448sw/rWzaOXJ3si4zSQYma0PVAkbyTTNZxnWyMsOP6D0tdy2/+qMHb1sou/k6z3O6hh621KOqbuvKrnuN3BDT1KuRhXvSlL8Qx7Z1b3y+FB+1V2HkP+djYalwfZosXlb/dc5CtJBz7IkCJAN74lXNQJvCjAGZvcDH1vovfdx17kBvS9jetfQpvQf1+PfJ5sSxP+v6lGv41zx9Njx3PQOsma/iZMsZZ75hNACBRS5AMrDIL9CeVr2wZW8cKUzNG9+TPSwIFNaAr7cXe3vXUBg0MP5xj98TphmmcakZ4FaaXtjf23phmjI5dq4EqkbkT7sy5VydlXIQyLsAyUDeI2CB2n/0WbcfFI8U/m95kFSYXhY+N7vHl091c2/vGvpjeTra2BFup/X3Nqe6ce8yiHAWqyouEFsuT5MYkU8ykMsIoNG7Q4BkYHeo5+icbZ0bWt2iN0bSuyeabRtN/rZ4WXzNdIPk2ruGwvTB5AyCI+r1IMyUdHy9+zBXYYp15GdaLt9PLzC+MiXXJT0mJSDQoADJQINQfG1mAiEJkEVnmRTGBFRGaIfpghf1rWt+Y6OltXcNbZDUWv7+z/p7WsLCP6k/VUvcTrOFcuqTUcCMBJLb61aWrp5RAXwZAQRmLEAyMGMyDphKYLIkQNLnvKnwzoFPHnrHTATbuwbvcteBFllYE+BX/T0tfzOT4yf7bmk52KjpRpkODavoNbob3lycmzKmFiAZIEIQWHgBkoGFN98jzxjWM4h2ROe7lXoCEh+7xb14dpgqOJuGt3cNbSktyxo+Zr/rX9f8uNmUU3tMzbiBe2ezk91c1IMydhVIzPbYEjaCmnqLXwQRQGAuBEgG5kIxx2WEnoBI0Tm7JgG6x6Rz42XxRdONC5iKr71reFjysSWCXV/o720JG+3MySe5u15lz/s5KTjHhYT1Fu7d746HXvWZQ/862Sp90/HUbhHd17sqDCLlgwAC8yhAMjCPuHty0W2dG9ois7NcVruhULhrXx/WD5iL9fSrpxbaQH9Pc/tcuVbtgsiMgtSsiV3wXqYoOq5/7aofzabQ8WTA9Jsl8ehRV/c+umrhqdmUyTEIIDC1AMlADiIk3LhDM2fbVV8hKo0HiKITzRU2aakM6ksI2i0unTvQ03zZXLHOZzJQ2rFtJLpGriOYUZD+ik2s3+D7plnVMbEFcgOb/aSvNyUggMD4DjBQ7KkC7V3D/ZK3lS/1jJ+s2zpve5Gs6YBI/pL6vQAluSGXnT2XSUDlesxnMhDO0dE19EGX/lXMKEj9TyCxeuBhxTg66oYLVlWtHtnoCSrTPknQGhXjewikF6BnIL3hoi0hLPnr0iVVFXS/dOloofOqz6z6S1gNsGjxgZW/mxX2cy/+Xfj/zQpHjSURU3zCjnSKP562x2GqU9RMLTy9v6clbPc6Z5/y5i7XyrWdGQXpWRObQH1+xeb7/6myQ2WjJSe3nZb8vP6e1rc1eizfQwCB2QuQDMzebtEfGV4PmEX9yYqWdwUM28n+RtIspunZLSa/KO3AwEbxJno2wqrBcftcJx6J99zHK9Ix/WtbftVo3fjergJHdw//fex+haQDTPqWRYVTr1t7aJgR0tDnud1/XFXwnd8Nu2C66ZUD61q+0tCBfAkBBFIJkAyk4lv8B7d3D/9Q7v+QqqZud5j8q/FYL8BgqrJmeHByBcK52Jeg3unDTo6bVq5cOrDuwPtnWD2+Xkego3v41e4eti1eJukvMn1oZMf2nhsvfPx904G1dQ2dYlJYbvquNK8apjsPf1/8Asef+ftlV3zyMQ8++x2bHnLjh1dOGzuLv0WLu4YkA4v7+sxJ7Tq6h3/p7k+aWWF+rxTd5PLPDPS0XDyzY+fu28nXBO7xGwd6V180d6XPf0nhhyycJV8/Zm5t3cMvN9enJT20rPygpMsiRR+4rufQW+tNOzyma+iwonSVpEe76ZoHl0Un/uD8Vdvn/ypxhsUm0N41+GnJTgvziStD2+brYWCxtX131YdkYHfJL/B527sGXyPZW2vW+a+txaBJ34s9vq24ZPna737yEZsWuJq7nC5ryUB71/CLJX+Hy54SKV7q0lJZVNqfOW8/Zh1rhh8VN/kXzBV6pqLExd1i8mvc9RWP/EYd2LpZf9p4eGR+oZv+VtKou71hoLf5S7s7/jj/wgq0rfntCkVL/8csel69M/f3tHDPmqdLAuw8wS7WYsP0wNDVf+SZf1q5dGRk5XUXtPx6sdY11Kutc8NNZtFTw/9dLBbffMP6w9Yu5vp2dA5udbPK03BppaTSk01IBkrTO1tz9m/Orb1z+FluOt/GNpxKJgWTXcpZDT5czHFB3SYEjj/Tlz24MyTNelosdZj8QHdbaqZ9JZV60ib7jDxk7/3y1cu2cJGTsx+mhYPlTHMj0NE1/EWXvyqUZtKH+npa/m1uSp67Uto6N5xqFr3A3Y8xs4dVlzzRzVn675HOipqKX77uE4fdNXc1yEZJIRE1i8ImVaeEAYKTJAbf8GXxKWlWrcyGxp5Vy/IiZHvF7tvNoldLOsQ9/qlZ4RlS/DiZbfXY9zezR5THkswCwNf297S+eRYHckgDAiQDDSDxld0n0N493C33tbLS0/UOV/zqgZ7Vc7ao0Uxb1nbqhuVaqlYVon8w6VmS1oyXEaoYe9iZMWQu2+Xaa7LyzewXfeuanzLT8+8p3w9rEmwvNh1mip5cenUV+2aZfae/p/mm2S5jvKfYLNZ2HPv6W/YZ2We/lZFFfx97/CRT9GiZnh0WICvPUppd1cNdKOTM5c94WWP/fb1J18Yeb5rrmUSzq+yeexTJwJ57bfeIloXXGU07RzZWniYssvf1rW1+10I3rqNz40tjxU8207lTndtlm+XFtQO9q89t6x56jrnCNLmpPpe66/MDvS3fWug2cT4EKgKl1UXHVhUN05HvHPvvFp7wwx25st5IGMxZN8Gt6f+qD1tz0y+V7OXkOXGEmX4Qx/HVUvRT/l0sXIySDCycNWeapUBH1/AWl4/tXCgN9ve0rJ5lUTM6rLxo08mSwv9M8vFNMvusXLdJ+ll/T8vPkl9MrpMwxckv7e9pecWMKseXEZilQHl58rbywmItkh6QNMPZRqWT3ynZ78oJw/4m/42bDbv7seZ2kEy3u/th5VcD9W78gya7zGP/Vv8FLdeFNSaieKSw0NOXZ8m4xx1GMrDHXdI9r0HtXUNhFcXkDfmI2pvuXLV6bF18nT7p6ovuD7j7NTJdVvDox40MwOzo3HiIK/5nNz0h7MBobo+Q6aCaOv/K5G/v62kNC/bwQWDOBMKYlsiibS7rrh/X9iMpPkyKSgtuuRevD70D7vEjqyrh2i5T2Hxq0Gzpsv6eQ347m0py05+N2vwfQzIw/8acIaVAW+fQi0z6ZhhBWP5sc9dHBnpb3pOy6NLhHV2Dx8uiV7h7mNYWRrzv8glPMLEXL1822vSNsJRzmvOWB9J9uF6Pg0k/ltmvYxU/oVj7L8b3pGE0+L17bV66ZOT+AgP90kTC3B9beeoPJe+6pLhvkutii6KHufQTj4u/Xrrt3p9c/YWnbJv7mlBi1gRIBrJ2xXJY3/JueOX3mGMA4b1i37qWMIBvxp/yngyvCD+W7nFbacBf/c+gTJd4HPfMR9dl+TXExyWVnsAme+9qZqWFltx9RLLHmqlqFcg4Lg5FZj+KZaOFWHdW9puYi0QiOUp8DD46z+RPHd/jzEweF98TxkjM+EJwQCqB0q6bo4WjFMcnSFZUadaN7TdRqA24Fz8XnuTnIhZSVZaDF70AycCiv0RUMAgc1Tl4XmT2lioNi17ev27V1xoVKu3VEBVeL/c3VY6pcwO+yj3+YcGjSxp5BdDouSf73lhCMNZ929AgrDoFTTeSOyQTIWEoH1pKJEyFKHL/YUgcKk+HbWfecahGRx9tip7uZm0mP37qsss1JiFIGwZTHh9u+hoJg/uio+QKC1WEAX37j28jbvZTuW9wj8OaIQPc+Of1cuyxhZMM7LGXds9rWGUw3sTUI9vucfHDsuh7Jv3VPQ4b4pRGRIf/HUWFTXFc3GYWhXnt4b/X/Zjp9jiOv6Co6eKBdav+sDvkwu6J8ugUxf4Qi+zoqepbW7/ZJhHj5bjvkFl4X7zLK5Lpy04sF6v4QwM9qxfdOhC743rO9pzJG3/kIUG0wxOxcE+53MvcbEBxcVDLdTOvamarzXFJAZIB4iEzAh3dG17tHv1PIxWe/iZW6nc/3+XfWIxPUuWpXqeGtkZRoSV2D7sA/qnymsDd93GPvxMGhsUev7AQWTGOtcOiaG/3kPhMs/309Ig/k+w+yfdXpBvk2uFxvN2iQlj3YeWkh7t2yPQbc90cR7pZcXzLYvSdvvkL841SEqjoqCjW4W4KN/7wP4mPX29mA7HbzfLizfPxumphWspZFrsAycBiv0LUr0qgvWvodknTTi2cIhn4lbsunavBh4v98pQHlA2aRQe4h3fLE+MNLCrs7XHxwNC1bBbdJ7enuIp9oU1T3XQ6zhx+lI36MbH0AZkqUz6noQjvr31AUXz5wLrVNy92t/moXynBs3JXv1mb5OHGH7r7K58huW4200DscbjxD8xHPSgTgXoCJAPEReYE2rqHrjZX3Y1MqhtjGyW/Qma3hJveQo0DyBzoLCtcGoNhUf8sDt9qrsu8YJf1r22+fBbHL/pDwhO/WaHFix42YDrcFW7+VTf+0Iah4BBHdrPiYnjXv6Dbgy96RCq4oAIkAwvKzcnmUiAMvosVHyIVDjH5fma2NHSRh3nSDKSaS+nJy2pwUaWGeg3MooH+nlXh2mXmU3XTD8vyjnX1h/EpySf+cnv8enO7OQ7v++nyz8w1zktFSQbycqVpJwLzJFBKyjze5OYrzZr2ieRhrfo2yY6azSlNflksC0/JNxdV/Pl3ew77xWzKmYtjSl37kfYff8oPN3rz/WsG9tU5VfnGP/bUz9S+ubgYlDGvAiQD88pL4QjkWyA8OUcetbmrTVaa5ZHY3nnGNoOSDZrirXLbGofxD5G2yrXVItuq2LaGEt13Dk3W5T4+Wj8ON/gmd8WtMu0fFngK/9vGRu8vl3z5roP5autrt0jx1tLTfvmmz+j+GV9TDlgkAiQDi+RCUA0E8iAwNnq+cLiVeg68dba9B/Np5dJ2k+6WQnIxdrOXaTC2kHAUQ49FeNLn/f58XgTKXnABkoEFJ+eECCCQFCg9re9ILKKT+KMv8aKNWiGy6PDKE7yXu+rHvlZ6kp9Nb8M9kpdmNZSm7oXXEh56GYpbGcVPfOZRgGQgj1edNiOwBwqUeh1Cd/9kn/BKYakGWaRnD7z4NCm1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUgv8fwA4c3jZPFf8AAAAAElFTkSuQmCC", "x": 10, "y": 230, "pages": "0-" } ] } ``` ## `Example` Response (B) To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/03c5c55183c74f8d94a4ec952e4e32ad/f1040-form-filled.pdf", "pageCount": 3, "error": false, "status": 200, "name": "f1040-form-filled", "remainingCredits": 60822 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "async": false, "inline": true, "name": "newDocument", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf", "annotations": [ { "text": "Sample Text 1", "x": 150, "y": 100, "size": 20, "pages": "0-" }, { "text": "sample text that is centered (can also set right or left alignment) ", "x": "10", "y": "10", "width": "500", "height": "200", "size": "7", "pages": "0", "alignment": "center" }, { "text": "Sample Text 2 - Click here to test link\r\n(CLICK ME!)", "x": 250, "y": 240, "size": 24, "pages": "0-", "color": "CCBBAA", "link": "https://pdf.co/", "fontName": "Comic Sans MS", "fontItalic": true, "fontBold": true, "fontStrikeout": false, "fontUnderline": true }, { "text": "Simple text 3", "x": 100, "y": 230, "size": 12, "pages": "0-", "type": "Text" }, { "text": "sample text 3 - input text field", "x": 100, "y": 170, "size": 16, "pages": "0-", "type": "TextField", "id": "textfield1" }, { "x": 200, "y": 120, "size": 16, "pages": "0-", "type": "Checkbox", "id": "checkbox2" }, { "x": 200, "y": 140, "size": 16, "pages": "0-", "type": "CheckboxChecked", "id": "checkbox3" } ``` ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "async": false, "inline": true, "name": "f1040-form-filled", "url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true", "imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200", "fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.pdf"; // Text annotation params const Type = "annotation"; const X = 400; const Y = 600; const Text = "APPROVED"; const FontName = "Times New Roman"; const FontSize = 24; const Color = "FF0000"; // * Add Text * // Prepare request to `PDF Edit` API endpoint var queryPath = `/v1/pdf/edit/add`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFileUrl, password: Password, annotations:[ { pages: Pages, x: X, y: Y, text: Text, fontname: FontName, size: FontSize, color: Color } ] }); var reqOptions = { host: "api.pdf.co", method: "POST", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download the output file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file).on("close", () => { console.log(`Generated PDF file saved to '${DestinationFile}' file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.error(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} 'https://api.pdf.co/v1/pdf/edit/add', CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => '', CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => 'POST', CURLOPT_POSTFIELDS =>'{ "async": false, "encrypt": false, "inline": true, "name": "f1040-form-filled", "url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true", "imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200", "fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789" }', CURLOPT_HTTPHEADER => array( 'Content-Type: application/json', 'x-api-key: ' ), )); $response = curl_exec($curl); curl_close($curl); echo $response; ?> ``` ```csharp theme={null} using Newtonsoft.Json.Linq; using System; using System.Globalization; using System.IO; using System.Net; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*****************************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination PDF file name const string DestinationFile = @".\result.pdf"; // Text annotation params const string Type = "annotation"; const int X = 400; const int Y = 600; const string Text = "APPROVED"; const string FontName = "Times New Roman"; const float FontSize = 24; const string FontColor = "FF0000"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // * Add text annotation * // Prepare requests params as JSON // See documentation: https://developer.pdf.co string jsonPayload = $@"{{ ""name"": ""{Path.GetFileName(DestinationFile)}"", ""url"": ""{SourceFileUrl}"", ""password"": ""{Password}"", ""annotations"": [ {{ ""x"": {X}, ""y"": {Y}, ""text"": ""{Text}"", ""fontname"": ""{FontName}"", ""size"": ""{FontSize.ToString(CultureInfo.InvariantCulture)}"", ""color"": ""{FontColor}"", ""pages"": ""{Pages}"" }} ] }}"; ; try { // URL of "PDF Edit" endpoint string url = "https://api.pdf.co/v1/pdf/edit/add"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download generated PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } finally { webClient.Dispose(); } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination PDF file name final static Path ResultFile = Paths.get(".\\result.pdf"); // Text annotation params private final static String Type2 = "annotation"; private final static int X2 = 400; private final static int Y2 = 600; private final static String Text = "APPROVED"; private final static String FontName = "Times New Roman"; private final static float FontSize = 24; private final static String Color = "FF0000"; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // * Add text annotation * // Prepare URL for `PDF Edit` API call String query = "https://api.pdf.co/v1/pdf/edit/add"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{ \"name\": \"%s\", \"url\": \"%s\", \"password\": \"%s\", annotations:[{ \"pages\": \"%s\", \"text\": \"%s\", \"x\": \"%s\", \"y\": \"%s\", \"fontname\": \"%s\", \"size\": \"%s\", \"color\": \"%s\" }] }", ResultFile.getFileName(), SourceFileUrl, Password, Pages, Text, X2, Y2, FontName, FontSize, Color); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated output file String resultFileUrl = json.get("url").getAsString(); // Download the image file downloadFile(webClient, resultFileUrl, ResultFile); System.out.printf("Generated file saved to \"%s\" file.", ResultFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile.toFile()); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} 'https://api.pdf.co/v1/pdf/edit/add', CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => '', CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => 'POST', CURLOPT_POSTFIELDS =>'{ "async": false, "encrypt": false, "inline": true, "name": "f1040-form-filled", "url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true", "imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200", "fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789" }', CURLOPT_HTTPHEADER => array( 'Content-Type: application/json', 'x-api-key: ' ), )); $response = curl_exec($curl); curl_close($curl); echo $response; ?> ``` *** ## Create Fillable PDF Forms You can create fillable PDF forms by adding editable text boxes and checkboxes. By using the [annotations\[\]](/api/pdf-add#pdf-add-annotations) attribute and setting the type to textfield or checkbox you can create form elements to be placed on your PDF . ### `Example` Payload (Create Fillable PDF Forms) ```json theme={null} { "async": false, "inline": true, "name": "newDocument", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf", "annotations":[ { "text":"sample prefilled text", "x": 10, "y": 30, "size": 12, "pages": "0-", "type": "TextField", "id": "textfield1" }, { "x": 100, "y": 150, "size": 12, "pages": "0-", "type": "Checkbox", "id": "checkbox2" }, { "x": 100, "y": 170, "size": 12, "pages": "0-", "link": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png", "type": "CheckboxChecked", "id":"checkbox3" } ] } ``` ### `Example` Response (Create Fillable PDF Forms) ```json theme={null} { "url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/d5c6efa549194ffaacb2eedd318e0320/newDocument.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzECMaDJJV7qKrpnGUrZHrwSKBATR5rxVlQoU0zj3r4jyHPt7yj4HoCIBi65IbMRWVX8qZZtKL9YGUzP%2FcemlqVd4Vi5%2B80Sg%2BymqQtaQ8qSFqKA82JnV%2BNBDatIigZIZha%2BrQM3jSC%2FZhX1zxsfLLsaH3K5nBnkjT3gi%2FZnx%2FgqrlIhf3m2xRFaTlgHrBADlK9KKPIijSusD4BTIo%2FQ433xx%2FQEaGWdX0nu4NuiByyXNPsBCAI3im9LMUCujjqF79ocyLHA%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHCSWKUQ4T/20200716/us-west-2/s3/aws4_request&X-Amz-Date=20200716T092641Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=2aa88d39aaf4b5891e4cb42d5675a64486098558d7159b37b75252209bdd6a95", "pageCount": 1, "error": false, "status": 200, "name": "newDocument", "remainingCredits": 77762 } ``` ### CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "async": false, "inline": true, "name": "newDocument", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf", "annotations":[ { "text":"sample prefilled text", "x": 10, "y": 30, "size": 12, "pages": "0-", "type": "TextField", "id": "textfield1" }, { "x": 100, "y": 150, "size": 12, "pages": "0-", "type": "Checkbox", "id": "checkbox2" }, { "x": 100, "y": 170, "size": 12, "pages": "0-", "link": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png", "type": "CheckboxChecked", "id":"checkbox3" } ] }' ``` ## Fill PDF Forms You can fill existing form fields in a PDF after identifying the form field names. Once form fields are identified then the [fields\[\]](#:~:text=height%20parameters%20provided.-,fields,-array%5Bobject%5D) attribute should be used to populate the fields by fieldName . ### `Example` Payload (Fill PDF Forms) ```json theme={null} { "async": false, "inline": true, "name": "f1040-filled", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "fields": [ { "fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]", "pages": "1", "text": "True" }, { "fieldName": "topmostSubform[0].Page1[0].f1_02[0]", "pages": "1", "text": "John A." }, { "fieldName": "topmostSubform[0].Page1[0].f1_03[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]", "pages": "1", "text": "123456789" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B.", "fontName": "Arial", "size": 6, "fontBold": true, "fontItalic": true, "fontStrikeout": true, "fontUnderline": true }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B." }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]", "pages": "1", "text": "987654321" } ] ``` ### `Example` Response (Fill PDF Forms) ```json theme={null} { "hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "url": "https://pdf-temp-files.s3.amazonaws.com/cd15a09771554bed88d6419c1e2f2b16/f1040-filled.pdf", "pageCount": 3, "error": false, "status": 200, "name": "f1040-filled.pdf", "remainingCredits": 99999369, "credits": 63 } ``` ### CURL ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "async": false, "inline": true, "name": "f1040-filled", "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf", "fields": [ { "fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]", "pages": "1", "text": "True" }, { "fieldName": "topmostSubform[0].Page1[0].f1_02[0]", "pages": "1", "text": "John A." }, { "fieldName": "topmostSubform[0].Page1[0].f1_03[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]", "pages": "1", "text": "123456789" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B.", "fontName": "Arial", "size": 6, "fontBold": true, "fontItalic": true, "fontStrikeout": true, "fontUnderline": true }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B." }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]", "pages": "1", "text": "987654321" } ], "annotations":[ { "text":"Sample Filled with PDF.co API using /pdf/edit/add. Get fields from forms using /pdf/info/fields. This text is be added on the first (0) and the last (!0) pages.", "x": 400, "y": 10, "width": 200, "height": 500, "size": 12, "pages": "0-", "color": "FF0000", "link": "https://pdf.co" } ] }' ``` ### Code samples (For Fill PDF Forms) ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf"; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.pdf"; // Runs processing asynchronously. Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success). Must be one of: true, false. const async = false; // Form field data var fields = [ { "fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]", "pages": "1", "text": "True" }, { "fieldName": "topmostSubform[0].Page1[0].f1_02[0]", "pages": "1", "text": "John A." }, { "fieldName": "topmostSubform[0].Page1[0].f1_03[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]", "pages": "1", "text": "123456789" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B." }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", "pages": "1", "text": "Joan B." }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]", "pages": "1", "text": "Doe" }, { "fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]", "pages": "1", "text": "987654321" } ]; // * Fill forms * // Prepare request to `PDF Edit` API endpoint var queryPath = `/v1/pdf/edit/add`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), password: Password, url: SourceFileUrl, async: async, fields: fields }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download the PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file).on("close", () => { console.log(`Generated PDF file saved to '${DestinationFile}' file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.error(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "**************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" def main(args = None): fillPDFForm() def fillPDFForm(): """Fill PDF form using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co payload = "{\n \"async\": false,\n \"encrypt\": false,\n \"name\": \"f1040-filled\",\n \"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf\",\n \"fields\": [\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]\",\n \"pages\": \"1\",\n \"text\": \"True\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].f1_02[0]\",\n \"pages\": \"1\",\n \"text\": \"John A.\"\n }, \n {\n \"fieldName\": \"topmostSubform[0].Page1[0].f1_03[0]\",\n \"pages\": \"1\",\n \"text\": \"Doe\"\n }, \n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]\",\n \"pages\": \"1\",\n \"text\": \"123456789\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n \"pages\": \"1\",\n \"text\": \"Joan B.\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n \"pages\": \"1\",\n \"text\": \"Joan B.\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]\",\n \"pages\": \"1\",\n \"text\": \"Doe\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]\",\n \"pages\": \"1\",\n \"text\": \"987654321\"\n } \n\n\n\n ],\n \"annotations\":[\n {\n \"text\":\"Sample Filled with PDF.co API using /pdf/edit/add. Get fields from forms using /pdf/info/fields\",\n \"x\": 10,\n \"y\": 10,\n \"size\": 12,\n \"pages\": \"0-\",\n \"color\": \"FFCCCC\",\n \"link\": \"https://pdf.co\"\n }\n ], \n \"images\": [ \n ]\n}" # Prepare URL for 'Fill PDF' API request url = "{}/pdf/edit/add".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=payload, headers={"x-api-key": API_KEY, 'Content-Type': 'application/json'}) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.Net; using System.Runtime.InteropServices; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*************************"; // Direct URL of source PDF file. const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf"; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // File name for generated output. Must be a String const string FileName = "f1040-form-filled"; // Destination File Name const string DestinationFile = "./result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Values to fill out pdf fields with built-in pdf form filler var fields = new List { new { fieldName = "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]", pages = "1", text = "True" }, new { fieldName = "topmostSubform[0].Page1[0].f1_02[0]", pages = "1", text = "John A." }, new { fieldName = "topmostSubform[0].Page1[0].f1_03[0]", pages = "1", text = "Doe" }, new { fieldName = "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]", pages = "1", text = "123456789" }, new { fieldName = "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]", pages = "1", text = "John B." }, new { fieldName = "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]", pages = "1", text = "Doe" }, new { fieldName = "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]", pages = "1", text = "987654321" } }; // If enabled, Runs processing asynchronously. Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, var async = false; // Prepare requests params as JSON // See documentation: https://developer.pdf.co Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileUrl); parameters.Add("name", FileName); parameters.Add("password", Password); parameters.Add("async", async); parameters.Add("fields", fields); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // URL of "PDF Edit" endpoint string url = "https://api.pdf.co/v1/pdf/edit/add"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download generated PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } finally { webClient.Dispose(); } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "****************************"; // Direct URL of source PDF file. final static String SourceFileUrl = "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf"; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination PDF file name final static Path ResultFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `PDF Edit` API call String query = "https://api.pdf.co/v1/pdf/edit/add"; // Prepare form filling data String fields = "[\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"True\"\n" + " },\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].f1_02[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"John A.\"\n" + " }, \n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].f1_03[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"Doe\"\n" + " }, \n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"123456789\"\n" + " },\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"Joan B.\"\n" + " },\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"Joan B.\"\n" + " },\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"Doe\"\n" + " },\n" + " {\n" + " \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]\",\n" + " \"pages\": \"1\",\n" + " \"text\": \"987654321\"\n" + " } \n" + " ]"; // Asynchronous Job String async = "false"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\n" + " \"url\": \"%s\",\n" + " \"async\": %s,\n" + " \"encrypt\": false,\n" + " \"name\": \"f1040-filled\",\n" + " \"fields\": %s"+ "}", SourceFileUrl, async, fields); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated output file String resultFileUrl = json.get("url").getAsString(); // Download the image file downloadFile(webClient, resultFileUrl, ResultFile); System.out.printf("Generated file saved to \"%s\" file.", ResultFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile.toFile()); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null}

Result:

" . $resultFileUrl . ""; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> ```
# Make Text Searchable Source: https://developer.pdf.co/api/pdf-change-text-searchable/searchable Convert scanned PDF or image files into text-searchable PDFs by running OCR and adding an invisible text layer for search and indexing. **Try it live:** [Make Text Searchable → API Tester](/api-tester/pdf-change-text-searchable/searchable) — send a real request from your browser. ## `POST /v1/pdf/makesearchable` This endpoint uses **300 DPI rendering** by default, which may significantly increase output file size-especially for image-heavy PDFs. To reduce file size, set a lower rendering resolution by using the [`profiles`](/api/profiles) attribute in the request body. For example, to set the rendering resolution to 72 DPI, use: ```json theme={null} { 'RenderingResolution': 72 } ``` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ---------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-make-searchable/sample.pdf", "lang": "eng", "pages": "", "name": "result.pdf", "password": "", "async": "false", "profiles": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/a0d52f35504e47148d1771fce875db7b/result.pdf", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 99033681, "credits": 35 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/makesearchable' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-make-searchable/sample.pdf", "lang": "eng", "pages": "", "name": "result.pdf", "password": "", "async": "false", "profiles": "" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more. const Language = "eng"; // Destination PDF file name const DestinationFile = "./result.pdf"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. MAKE UPLOADED PDF FILE SEARCHABLE makePdfSearchable(API_KEY, uploadedFileUrl, Password, Pages, Language, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function makePdfSearchable(apiKey, uploadedFileUrl, password, pages, language, destinationFile) { // Prepare request to `Make Searchable PDF` API endpoint var queryPath = `/v1/pdf/makesearchable`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, lang: language, url: uploadedFileUrl, async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url, destinationFile); } else { // Service reported error console.log("makePdfSearchable(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("makePdfSearchable(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(destinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${destinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more. Language = "eng" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): makeSearchablePDF(uploadedFileUrl, DestinationFile) def makeSearchablePDF(uploadedFileUrl, destinationFile): """Make Uploaded PDF file Searchable using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["lang"] = Language parameters["url"] = uploadedFileUrl # Prepare URL for 'Make Searchable PDF' API request url = "{}/pdf/makesearchable".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more. const string Language = "eng"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. MAKE UPLOADED PDF FILE SEARCHABLE // URL for `Make Searchable PDF` API call var url = "https://api.pdf.co/v1/pdf/makesearchable"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("url", uploadedFileUrl); parameters.Add("pages", Pages); parameters.Add("lang", Language); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more. final static String Language = "eng"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. MAKE UPLOADED PDF FILE SEARCHABLE MakePdfSearchable(webClient, API_KEY, DestinationFile, Password, Pages, Language, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void MakePdfSearchable(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String language, String uploadedFileUrl) throws IOException { // Prepare URL for `Make Searchable PDF` API call String query = "https://api.pdf.co/v1/pdf/makesearchable"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"lang\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, language, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Make PDF Searchable Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function MakePdfSearchable($apiKey, $uploadedFileUrl, $pages, $ocrLanguage) { // Prepare URL for `Make Searchable PDF` API call $url = "https://api.pdf.co/v1/pdf/makesearchable"; // Prepare requests params $parameters = array(); $parameters["name"] = "result.pdf"; $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; $parameters["lang"] = $ocrLanguage; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the result file echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# Make Text Unsearchable Source: https://developer.pdf.co/api/pdf-change-text-searchable/unsearchable Convert a PDF into a flat, image-only version that is no longer text-searchable, by rasterizing each page as a scanned image. **Try it live:** [Make Text Unsearchable → API Tester](/api-tester/pdf-change-text-searchable/unsearchable) — send a real request from your browser. ## `POST /v1/pdf/makeunsearchable` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ---------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "pages": "", "name": "result.pdf", "password": "", "async": "false", "profiles": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/6b755238963a472abf67fd5e7ffafd79/result.pdf", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 327244, "credits": 35 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/makeunsearchable' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "pages": "", "name": "result.pdf", "password": "", "async": "false", "profiles": "" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.pdf"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. MAKE UPLOADED PDF FILE UNSEARCHABLE makePdfUnSearchable(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function makePdfUnSearchable(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `Make UnSearchable PDF` API endpoint var queryPath = `/v1/pdf/makeunsearchable`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl, async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url, destinationFile); } else { // Service reported error console.log("makePdfUnSearchable(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("makePdfUnSearchable(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(destinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${destinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import requests import json # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "*****************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" fileName = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf" url = "{}/pdf/makeunsearchable?url={}".format(BASE_URL, fileName) # Execute request and get response as JSON response = requests.get(url, headers={"x-api-key": API_KEY}) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL of unsearchable PDF unsearchableFile = json["url"] print(unsearchableFile) ``` ```php theme={null} ``` # PDF Compress Source: https://developer.pdf.co/api/pdf-compress POST /v2/pdf/compress Compress a PDF with ready-to-use presets or advanced image and font controls. **Try it live:** [PDF Compress → API Tester](/api-tester/pdf-compress) — send a real request from your browser. ## `POST /v2/pdf/compress` This is the current PDF compression endpoint. The legacy PDF Optimize V1 endpoint is deprecated. ## Quick start A request containing only `url` uses the standard configuration — the same image optimization as Adobe Acrobat Pro — so you can start compressing immediately. Start with the `medium` preset for balanced compression. Choose another preset only when you need lighter or stronger compression. To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
## Choose a preset For most files, choose one of four `compression_level` values. You do not need to understand the advanced configuration to use them. * `low` — Light compression with the highest image resolution. * `medium` — Balanced compression based on Adobe Standard resolution targets. * `high` — Stronger compression with lower image resolution. * `aggressive` — The smallest images and strongest compression of the four presets. You can also set `color_quality` from `1` to `100`. Higher values preserve more color and grayscale image quality and usually produce larger files. The default is `80`. ```json Preset request body theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", "compression_level": "high", "color_quality": 80 } ``` Exact PPI thresholds used by each preset. | Preset | Color and grayscale | Monochrome | | ------------ | -------------------------- | -------------------------- | | `low` | 200 PPI when above 300 PPI | 400 PPI when above 600 PPI | | `medium` | 150 PPI when above 225 PPI | 300 PPI when above 450 PPI | | `high` | 127 PPI when above 172 PPI | 180 PPI when above 270 PPI | | `aggressive` | 96 PPI when above 120 PPI | 100 PPI when above 150 PPI | ## Request body Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` There are no query parameters. URL of the source PDF. The endpoint processes the complete document. The URL must be reachable by PDF.co — see [supported file sources](/api/url-input-and-request-limits#supported-file-sources). For a protected source location, provide a temporary or presigned URL. Ready-to-use compression profile: `low`, `medium`, `high`, or `aggressive`. When omitted, the standard configuration is applied. Presets also enable font subsetting and font stream compression. JPEG2000 quality for color and grayscale images, from `1` for the smallest file and lowest quality to `100` for the highest quality. It can be used with or without `compression_level`. Password for opening an encrypted source PDF. Omit it for an unprotected PDF. Set to `true` for large or long-running documents. The initial response includes `jobId`; use the [Background Job Check endpoint](/api/job-check) to retrieve the final status. Also see [Webhooks & Callbacks](/api/webhooks). Callback URL notified when an asynchronous job finishes. Use it with `async: true`. Output file name. The endpoint appends `.pdf` when needed. If omitted, it derives the name from `url` when possible. Number of minutes before the temporary output URL expires. After this period, generated files are deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum retention depends on your subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). The current V2 Compress implementation does not use `httpusername` or `httppassword`. The old `profiles.outputDataFormat` and `profiles.JPEGQuality` descriptions also do not apply to this endpoint.
## Advanced configuration Use `config` only when a preset is not enough. It is a partial object: you only send the values you want to override. 1. PDF.co starts with the standard configuration. 2. It applies the selected `compression_level`, when supplied. 3. It applies `color_quality`, when supplied. 4. It deep-merges `config` as the final override. | Request | Image resolution | Color and grayscale encoding | Fonts | | ------------------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------- | ------------------- | | `url` only | Standard targets | JPEG, quality 60 | Unchanged | | `compression_level` | Targets from the selected preset | JPEG2000, quality 80 unless overridden | Subset and compress | | `color_quality` only | Standard targets | JPEG2000 at the selected quality | Unchanged | | `config` only | Start with the standard configuration, then replace only the values supplied in `config`. | | | | Preset or quality plus `config` | Apply the preset and quality first, then replace only the values supplied in `config`. | | | Specify the narrowest override you need. Every omitted value continues to come from the base configuration, or from the preset and `color_quality` you selected. Each preset resolves to an effective configuration for the first compression pass. All presets use JPEG2000 quality 80 for color and grayscale images, CCITT Group 4 for monochrome images, font subsetting and compression, and garbage collection level 4. Their resolution targets differ per the **Preset resolution targets** table above. To represent another preset, change the four PPI values using that table. A `color_quality` of 80 maps to JPEG2000 `rates` with `quality_layers: [20]`. The effective values produced by the `medium` preset: ```json theme={null} { "config": { "images": { "color": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 }, "compression": { "skip": false, "compression_format": "jpeg2000", "compression_params": { "quality_mode": "rates", "quality_layers": [20] } } }, "grayscale": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 }, "compression": { "skip": false, "compression_format": "jpeg2000", "compression_params": { "quality_mode": "rates", "quality_layers": [20] } } }, "monochrome": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 300, "threshold_ppi": 450 }, "compression": { "skip": false, "compression_format": "ccitt_g4", "compression_params": {} } } }, "fonts": { "subset": true, "compress": true }, "save": { "garbage": 4 } } } ``` Use `compression_level` when a preset already fits your needs. The configuration above produces the same first pass, but explicit `config` values are preserved during the standard fallback retry, so manually copying a complete preset can change fallback behavior. ### Common config recipes These examples show only the request fields relevant to the change. Add the same fields to the request body together with `url`. Keep the base encoding and fonts, but retain more color and grayscale detail. Images are reduced to 200 PPI only when their effective resolution is above 300 PPI. Monochrome images keep the base settings. ```json theme={null} { "config": { "images": { "color": { "downsample": { "downsample_ppi": 200, "threshold_ppi": 300 } }, "grayscale": { "downsample": { "downsample_ppi": 200, "threshold_ppi": 300 } } } } } ``` Preserve pixel dimensions while still applying JPEG2000 compression. ```json theme={null} { "config": { "images": { "color": { "downsample": { "skip": true } }, "grayscale": { "downsample": { "skip": true } } } } } ``` Do not downsample or re-encode monochrome images. ```json theme={null} { "config": { "images": { "monochrome": { "skip": true } } } } ``` Start with `high`, preserve more image quality, and make two exceptions. This keeps the `high` resolution targets, uses color quality 90, leaves monochrome images unchanged, and disables font subsetting. Font stream compression remains enabled. ```json theme={null} { "compression_level": "high", "color_quality": 90, "config": { "images": { "monochrome": { "skip": true } }, "fonts": { "subset": false } } } ``` ### What each skip setting does | Setting | What it skips | What can still run | | -------------------- | ------------------------------------------------------ | ------------------------------------------------------------ | | `images..skip` | All optimization for that image type | Other image types and fonts | | `downsample.skip` | Resolution reduction | Image re-encoding | | `compression.skip` | The selected JPEG, JPEG2000, CCITT, or ZIP compression | Downsampling; resized image data may still be written as PNG | Advanced image, font, and save controls. Settings for color, grayscale, and monochrome images. `color`, `grayscale` & `monochrome` all use the same object schema: Skip both downsampling and re-encoding for this image type. Control resolution reduction. Preserve image dimensions while allowing re-encoding. Target resolution when the image is above `threshold_ppi`. Default is `150` for `color` and `grayscale`, `300` for `monochrome`. Minimum effective resolution that triggers downsampling. Default is `225` for `color` and `grayscale`, `450` for `monochrome`. Control image re-encoding. Skip the selected compression format while allowing downsampling. Resized image data may still be written as PNG. `jpeg`, `jpeg2000`, `ccitt_g4`, `ccitt_g3`, or `zip`. Use CCITT for monochrome images. JPEG or JPEG2000 quality settings. JPEG quality from `1` to `100`. This field is used only when `compression_format` is `jpeg`. JPEG2000 mode: `rates` or `dB`. JPEG2000 quality layers. When supplied directly without valid layers, the fallback is `[30]` for `rates` or `[38.0, 34.0, 30.0]` for `dB`. Same object schema as `color`. Same object schema as `color`. Font subsetting and stream compression. Remove unused glyphs from fonts that can be subset. Compress font streams while saving the PDF. PDF cleanup settings. Garbage collection level from `0` to `4`. * `0` — none * `1` — remove unused objects * `2` — compact xref * `3` — merge duplicate objects * `4` — detect duplicate stream content Fallback base used when the first compression pass is not smaller. This is also the configuration applied when the request contains only `url`. ```json theme={null} { "images": { "color": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 }, "compression": { "skip": false, "compression_format": "jpeg", "compression_params": { "quality": 60 } } }, "grayscale": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 }, "compression": { "skip": false, "compression_format": "jpeg", "compression_params": { "quality": 60 } } }, "monochrome": { "skip": false, "downsample": { "skip": false, "downsample_ppi": 300, "threshold_ppi": 450 }, "compression": { "skip": false, "compression_format": "ccitt_g4", "compression_params": {} } } }, "fonts": { "subset": false, "compress": false }, "save": { "garbage": 4 } } ``` ## Behavior notes * Compression results depend on the source PDF. Presets do not promise a fixed reduction, and two levels can produce the same file size. * When the first compression pass does not produce a smaller PDF, PDF.co retries with the standard configuration while preserving explicit `config` overrides. If the retry is also not smaller, it returns the original PDF. * Each re-encoded image is kept only when its new stream is smaller. If effective PPI cannot be determined, downsampling is skipped but re-encoding can still run at the original dimensions. ## Responses A synchronous success returns the final temporary output URL. An asynchronous request returns a `jobId` and a reserved URL that should be used only after the job succeeds — poll it via [Background Job Check](/api/job-check). Errors return `error: true` with a status code and message — see the response examples and [Response Codes](/api/response-codes). Number of pages in the output PDF. `false` for a successful request. PDF.co status code. Success returns `200`. For more information, see [Response Codes](/api/response-codes). Credits consumed by the request. Credits remaining for the account. Processing duration in milliseconds. Temporary URL of the result. With size protection, it can point to a copy of the original PDF. Output file name. UTC timestamp when the temporary URL expires. Present in the initial response when `async` is `true`. **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
**Sample request** ```bash cURL (minimal) theme={null} curl --request POST \ --url 'https://api.pdf.co/v2/pdf/compress' \ --header 'Content-Type: application/json' \ --header 'x-api-key: YOUR_API_KEY' \ --data '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf" }' ``` ```bash cURL (preset) theme={null} curl --request POST \ --url 'https://api.pdf.co/v2/pdf/compress' \ --header 'Content-Type: application/json' \ --header 'x-api-key: YOUR_API_KEY' \ --data '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", "compression_level": "high", "color_quality": 80 }' ``` ```bash cURL (config override) theme={null} curl --request POST \ --url 'https://api.pdf.co/v2/pdf/compress' \ --header 'Content-Type: application/json' \ --header 'x-api-key: YOUR_API_KEY' \ --data '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", "compression_level": "high", "config": { "images": { "monochrome": { "compression": { "compression_format": "ccitt_g4" } } }, "fonts": { "subset": false } } }' ``` ```javascript Node.js theme={null} const response = await fetch( "https://api.pdf.co/v2/pdf/compress", { method: "POST", headers: { "Content-Type": "application/json", "x-api-key": "YOUR_API_KEY", }, body: JSON.stringify({ url: "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", compression_level: "medium", color_quality: 80, name: "compressed.pdf", }), } ); const result = await response.json(); if (!response.ok || result.error) { throw new Error(result.message ?? "PDF.co request failed"); } console.log(result.url); ``` ```python Python theme={null} import requests response = requests.post( "https://api.pdf.co/v2/pdf/compress", headers={"x-api-key": "YOUR_API_KEY"}, json={ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", "compression_level": "medium", "color_quality": 80, "name": "compressed.pdf", }, ) result = response.json() if result.get("error"): raise RuntimeError(result.get("message", "PDF.co request failed")) print(result["url"]) ``` ```csharp C# theme={null} using System.Collections.Generic; using System.Net.Http; using System.Text; using System.Text.Json; var client = new HttpClient(); client.DefaultRequestHeaders.Add("x-api-key", "YOUR_API_KEY"); var payload = JsonSerializer.Serialize(new Dictionary { ["url"] = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", ["compression_level"] = "medium", ["color_quality"] = 80, ["name"] = "compressed.pdf", }); var response = await client.PostAsync( "https://api.pdf.co/v2/pdf/compress", new StringContent(payload, Encoding.UTF8, "application/json") ); var json = await response.Content.ReadAsStringAsync(); Console.WriteLine(json); ``` ```java Java theme={null} OkHttpClient client = new OkHttpClient(); String jsonPayload = "{\"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf\", \"compression_level\": \"medium\", \"color_quality\": 80, \"name\": \"compressed.pdf\"}"; Request request = new Request.Builder() .url("https://api.pdf.co/v2/pdf/compress") .addHeader("x-api-key", "YOUR_API_KEY") .addHeader("Content-Type", "application/json") .post(RequestBody.create(MediaType.parse("application/json"), jsonPayload)) .build(); try (Response response = client.newCall(request).execute()) { System.out.println(response.body().string()); } ``` ```php PHP theme={null} "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf", "compression_level" => "medium", "color_quality" => 80, "name" => "compressed.pdf", ]); $curl = curl_init("https://api.pdf.co/v2/pdf/compress"); curl_setopt_array($curl, [ CURLOPT_HTTPHEADER => [ "x-api-key: YOUR_API_KEY", "Content-Type: application/json", ], CURLOPT_POST => true, CURLOPT_RETURNTRANSFER => true, CURLOPT_POSTFIELDS => $payload, ]); echo curl_exec($curl); curl_close($curl); ``` ```json 200 theme={null} { "pageCount": 2, "error": false, "status": 200, "credits": 70, "remainingCredits": 999860, "duration": 8768, "url": "https://pdf-temp-files.s3.amazonaws.com/example/sample.pdf", "name": "sample.pdf", "outputLinkValidTill": "2026-08-08T12:00:00+00:00" } ``` ```json 400 theme={null} { "error": true, "status": 400, "message": "Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password)." } ``` ```json 401 theme={null} { "error": true, "status": 401, "message": "Unauthorized. Authentication is required and has failed or has not yet been provided." } ``` ```json 402 theme={null} { "error": true, "status": 402, "message": "Not enough credits." } ``` ```json 403 theme={null} { "error": true, "status": 403, "message": "Access forbidden for input URL." } ``` ```json 404 theme={null} { "error": true, "status": 404, "message": "The requested resource could not be found." } ``` ```json 408 theme={null} { "error": true, "status": 408, "message": "The server timed out waiting for the request." } ``` ```json 429 theme={null} { "error": true, "status": 429, "message": "Too many requests in a given time period." } ``` ```json 441 theme={null} { "error": true, "status": 441, "message": "Invalid Password. Password protected document." } ``` ```json 442 theme={null} { "error": true, "status": 442, "message": "Input document is damaged or of incorrect type." } ``` ```json 443 theme={null} { "error": true, "status": 443, "message": "Permissions. The operation is prohibited by document security settings." } ``` ```json 444 theme={null} { "error": true, "status": 444, "message": "Profiles parsing error. Please ensure that the configuration is supported." } ``` ```json 445 theme={null} { "error": true, "status": 445, "message": "Timeout error. For large documents, use asynchronous mode (async=true) and check status via /job/check." } ``` ```json 446 theme={null} { "error": true, "status": 446, "message": "Some files required for conversion are missing." } ``` ```json 447 theme={null} { "error": true, "status": 447, "message": "Invalid template." } ``` ```json 448 theme={null} { "error": true, "status": 448, "message": "Invalid URL or HTML. Ensure the provided URL is valid and accessible." } ``` ```json 449 theme={null} { "error": true, "status": 449, "message": "Invalid index range. Page index is out of range." } ``` ```json 450 theme={null} { "error": true, "status": 450, "message": "Invalid page range specified." } ``` ```json 452 theme={null} { "error": true, "status": 452, "message": "Invalid URL." } ``` ```json 454 theme={null} { "error": true, "status": 454, "message": "Invalid parameters." } ``` ```json 500 theme={null} { "error": true, "status": 500, "message": "Something went wrong. Please try again or contact support." } ``` # PDF Delete Pages Source: https://developer.pdf.co/api/pdf-delete-pages Deletes selected pages inside a PDF file. **Try it live:** [PDF Delete Pages → API Tester](/api-tester/pdf-delete-pages) — send a real request from your browser. ## `POST /v1/pdf/edit/delete-pages` The `pages` parameter is 1-based, meaning the first page is `1` and not `0`. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *Yes* | - | Specify pages as comma-separated page numbers and ranges to delete (e.g. "1, 2, 5-10" or "3-" for page 3 to the end). The first-page index is 1. Inverted page numbers (e.g. "!1" for the last page) are not supported by this endpoint. This parameter is required: omitting it, or sending an empty or whitespace-only value, returns HTTP 400. The input must be in string format. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf", "pages": "1-2", "name": "result.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/d15e5b2c89c04484ae6ac7244ac43ac2/result.pdf", "pageCount": 2, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 60100 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/delete-pages' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf", "pages": "1-2", "name": "result.pdf", "async": false }' ``` ```javascript theme={null} // `request` module is required for file upload. // Use "npm install request" command to install. var request = require('request'); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/pdf/edit/delete-pages', 'headers': { 'x-api-key': '{{x-api-key}}' }, formData: { 'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf', 'name': 'result.pdf', 'pages': '1-2' } }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ``` ```python theme={null} import requests url = "https://api.pdf.co/v1/pdf/edit/delete-pages" # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ payload = {'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf', 'name': 'result.pdf', 'pages': '1-2'} files = [ ] headers = { 'x-api-key': '{{x-api-key}}' } response = requests.request("POST", url, headers=headers, json = payload, files = files) print(response.text.encode('utf8')) ``` ```csharp theme={null} using System; using RestSharp; namespace HelloWorldApplication { class HelloWorld { static void Main(string[] args) { var client = new RestClient("https://api.pdf.co/v1/pdf/edit/delete-pages"); client.Timeout = -1; var request = new RestRequest(Method.POST); request.AddHeader("x-api-key", "{{x-api-key}}"); request.AlwaysMultipartFormData = true; // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ request.AddParameter("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"); request.AddParameter("name", "result.pdf"); request.AddParameter("pages", "1-2"); IRestResponse response = client.Execute(request); Console.WriteLine(response.Content); } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("text/plain"); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM) .addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf") .addFormDataPart("name", "result.pdf") .addFormDataPart("pages", "1-2") .build(); Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/edit/delete-pages") .method("POST", body) .addHeader("x-api-key", "{{x-api-key}}") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} "https://api.pdf.co/v1/pdf/edit/delete-pages", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "POST", CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf','name' => 'result.pdf','pages' => '1-2'), CURLOPT_HTTPHEADER => array( "x-api-key: {{x-api-key}}" ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; ```
# Extract Attachment Source: https://developer.pdf.co/api/pdf-extract-attachments Extracts attachments from a PDF file. **Try it live:** [Extract Attachment → API Tester](/api-tester/pdf-extract-attachments) — send a real request from your browser. ## `POST /v1/pdf/attachments/extract` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `pageCount` | integer | Number of pages in the PDF document. | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-attachments/attachments.pdf", "inline": false, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/DO1TAIHEZR5P9QLI7ICYM9DI0AAH57HY/sample.png", "https://pdf-temp-files.s3.amazonaws.com/EOINIMD7X48JSOB1G8ETLVPOFZLM1NJ2/SampleMetafile.emf", "https://pdf-temp-files.s3.amazonaws.com/3LW4BXNSPAE0WQTG5DPMXX498OCPNU4Q/ab.tif" ], "pageCount": 3, "error": false, "status": 200, "name": "attachments.json", "credits": 24, "duration": 1211, "remainingCredits": 98003902 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/attachments/extract' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-attachments/attachments.pdf", "inline": false, "async": false }' ``` ```javascript theme={null} var https = require("https"); var fs = require("fs"); var path = require("path"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf"; // Prepare request for API endpoint var queryPath = `/v1/pdf/attachments/extract`; // JSON payload for api request var jsonPayload = JSON.stringify({ url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { let responseData = ''; response.setEncoding("utf8"); response.on("data", (chunk) => { responseData += chunk; }); response.on("end", () => { // Parse JSON response var data = JSON.parse(responseData); if (data.error == false) { // Download extracted files data.urls.forEach((url) => { var localFileName = path.basename(url); var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated file saved as "${localFileName}" file.`); }); }); }, this); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.error(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import requests import json url = "https://api.pdf.co/v1/pdf/attachments/extract" payload = json.dumps({ "url": "https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf", "inline": True, "async": False }) headers = { 'Content-Type': 'application/json', 'x-api-key': '__Replace_With_Your_PDFco_API_Key__' } response = requests.request("POST", url, headers=headers, data=payload) print(response.text) ``` ```csharp theme={null} using System; using RestSharp; namespace HelloWorldApplication { class HelloWorld { static void Main(string[] args) { var client = new RestClient("https://api.pdf.co/v1/pdf/attachments/extract"); client.Timeout = -1; var request = new RestRequest(Method.POST); request.AddHeader("Content-Type", "application/json"); request.AddHeader("x-api-key", "__Replace_With_Your_PDFco_API_Key__"); var body = @"{" + "\n" + @" ""url"": ""https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf""," + "\n" + @" ""inline"": true," + "\n" + @" ""async"": false" + "\n" + @"}"; request.AddParameter("application/json", body, ParameterType.RequestBody); IRestResponse response = client.Execute(request); Console.WriteLine(response.Content); } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("application/json"); RequestBody body = RequestBody.create(mediaType, "{\n \"url\": \"https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf\",\n \"inline\": true,\n \"async\": false\n}"); Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/attachments/extract") .method("POST", body) .addHeader("Content-Type", "application/json") .addHeader("x-api-key", "__Replace_With_Your_PDFco_API_Key__") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} Cloud API asynchronous "Extract PDF Attachment" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo ""; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF Find Text Source: https://developer.pdf.co/api/pdf-find/basic Find text in PDF and get coordinates. Supports regular expressions. **Try it live:** [PDF Find Text → API Tester](/api-tester/pdf-find/basic) — send a real request from your browser. ## `POST /v1/pdf/find` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`. | Attribute | Type | Required | Default | Description | | ------------------------------------------------------ | ------- | -------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `searchString` | string | *Yes* | - | Text to search can support regular expressions if you set the `regexSearch` param to true. | | `wordMatchingMode` | string | *No* | None | WordMatchingMode defines how search terms match PDF text. Modes: `None` (exact string match only), `SmartMatch` (default; flexible word boundary match, includes letters/digits/punctuation), `ExactMatch` (strict word boundaries, whole-word match only). | | `regexSearch` | boolean | *No* | `false` | Set to true to enable regular expression search for the `searchString(s)` parameter. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `ColumnDetectionMode` | string | *No* | ContentGroupsAndBorders | Controls column detection/alignment in PDF table extraction. Modes: `ContentGroupsAndBorders` (default; text + lines), `ContentGroups` (text grouping only), `Borders` (lines only), `BorderedTables` (OCR-based for bordered tables), `ContentGroupsAI` (AI for dense/complex layouts). | |     `DetectionMinNumberOfRows` | integer | *No* | 1 | Minimum number of rows to detect in a table | |     `DetectionMinNumberOfColumns` | integer | *No* | 1 | Minimum number of columns to detect in a table | |     `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | *No* | `0` | Maximum number of invalid subsequent rows allowed in a table | |     `DetectionMinNumberOfLineBreaksBetweenTables` | integer | *No* | `0` | Minimum number of line breaks between tables | |     `EnhanceTableBorders` | boolean | *No* | `true` | Enhance table borders or not | |     `OCRDetectPageRotation` | boolean | *No* | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See [Support page rotation](#support-page-rotation) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | | `requestParametersDocument` | string | *No* | - | | | `responseParameters` | object | *No* | - | - | |     `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) | |     `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | |     `message` | string | *No* | - | Message of the request | |     `credits` | integer | *No* | - | Number of credits consumed by the request | |     `remainingCredits` | integer | *No* | - | Number of credits remaining in the account | |     `duration` | integer | *No* | - | Time taken for the operation in milliseconds | |     `errorCode` | integer | *No* | - | Error code of the request (400, 401, 402, 403, 404, 500, etc.) | ### Support page rotation This endpoint supports **PDF** page rotation as follows: ```json theme={null} { "profiles": "{ 'OCRDetectPageRotation': true }" } ``` ### Find only bordered tables You can limit search to bordered tables only by enabling the *legacy table* search mode with the following `profiles` config: ```json theme={null} { "profiles": "{ 'Mode': 'Legacy', 'ColumnDetectionMode': 'BorderedTables', 'DetectionMinNumberOfRows': 1, 'DetectionMinNumberOfColumns': 1, 'DetectionMaxNumberOfInvalidSubsequentRowsAllowed': 0, 'DetectionMinNumberOfLineBreaksBetweenTables': 0, 'EnhanceTableBorders': false }" } ``` ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "async": "false", "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "searchString": "Invoice Date \\d+/\\d+/\\d+", "regexSearch": "true", "name": "output", "pages": "0-", "inline": "true", "wordMatchingMode": "", "password": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": [ { "text": "Invoice Date 01/01/2016", "left": 436.5400085449219, "top": 130.4599995137751, "width": 122.85311957550027, "height": 11.040000486224898, "pageIndex": 0, "bounds": { "location": { "isEmpty": false, "x": 436.54, "y": 130.46 }, "size": "122.853119, 11.0400009", "x": 436.54, "y": 130.46, "width": 122.853119, "height": 11.0400009, "left": 436.54, "top": 130.46, "right": 559.3931, "bottom": 141.5, "isEmpty": false }, "elementCount": 1, "elements": [ { "index": 0, "left": 436.5400085449219, "top": 130.4599995137751, "width": 122.85311957550027, "height": 11.040000486224898, "angle": 0, "text": "Invoice Date 01/01/2016", "isNewLine": true, "fontIsBold": true, "fontIsItalic": false, "fontName": "Helvetica-Bold", "fontSize": 11, "fontColor": "0, 0, 0", "fontColorAsOleColor": 0, "fontColorAsHtmlColor": "#000000", "bounds": { "location": { "isEmpty": false, "x": 436.54, "y": 130.46 }, "size": "122.853119, 11.0400009", "x": 436.54, "y": 130.46, "width": 122.853119, "height": 11.0400009, "left": 436.54, "top": 130.46, "right": 559.3931, "bottom": 141.5, "isEmpty": false } } ] } ], "pageCount": 1, "error": false, "status": 200, "name": "output", "remainingCredits": 59970 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/find' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "async": "false", "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "searchString": "Invoice Date \\d+/\\d+/\\d+", "regexSearch": "true", "name": "output", "pages": "0-", "inline": "true", "wordMatchingMode": "", "password": "" }' ``` ```javascript theme={null} // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Search string. const SearchString = '[4-9][0-9].[0-9][0-9]'; // Regular expression to find numbers in format dd.dd and between 40.00 to 99.99 // Enable regular expressions (Regex) const RegexSearch = 'True'; // Prepare URL for PDF text search API call. // See documentation: https://developer.pdf.co var query = `https://api.pdf.co/v1/pdf/find`; let reqOptions = { uri: query, headers: { "x-api-key": API_KEY }, formData: { password: Password, pages: Pages, url: SourceFileUrl, searchString: SearchString, regexSearch: RegexSearch } }; // Send request request.post(reqOptions, function (error, response, body) { if (error) { return console.error("Error: ", error); } // Parse JSON response let data = JSON.parse(body); for (let index = 0; index < data.body.length; index++) { const element = data.body[index]; console.log("Found text " + element["text"] + " at coordinates " + element["left"] + ", " + element["top"]); } }); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Search string. SearchString = "\d{1,}\.\d\d" # Regular expression to find numbers like '100.00' # Note: do not use `+` char in regex, but use `{1,}` instead. # `+` char is valid for URL and will not be escaped, and it will become a space char on the server side. # Enable regular expressions (Regex) RegexSearch = True def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): searchTextInPDF(uploadedFileUrl) def searchTextInPDF(uploadedFileUrl): """Search Text using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {} parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl parameters["searchString"] = SearchString parameters["regexSearch"] = RegexSearch # Prepare URL for 'PDF Text Search' API request url = "{}/pdf/find".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Display found information for item in json["body"]: print(f"Found text {item['text']} at coordinates {item['left']}, {item['top']}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Search string. const string SearchString = @"\d{1,}\.\d\d"; // Regular expression to find numbers like '100.00' // Note: do not use `+` char in regex, but use `{1,}` instead. // `+` char is valid for URL and will not be escaped, and it will become a space char on the server side. // Enable regular expressions (Regex) const bool RegexSearch = true; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. MAKE UPLOADED PDF FILE SEARCHABLE // URL for `PDF Text Search` API call // See documentation: https://developer.pdf.co string url = "https://api.pdf.co/v1/pdf/find"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); parameters.Add("searchString", SearchString); parameters.Add("regexSearch", RegexSearch); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { foreach (JToken item in json["body"]) { Console.WriteLine($"Found text \"{item["text"]}\" at coordinates {item["left"]}, {item["top"]}"); } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException ex) { Console.WriteLine(ex.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Search string. final static String SearchString = "\\d{1,}\\.\\d\\d"; // Regular expression to find numbers like '100.00' // Note: do not use `+` char in regex, but use `{1,}` instead. // `+` char is valid for URL and will not be escaped, and it will become a space char on the server side. // Enable regular expressions (Regex) final static boolean RegexSearch = true; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for PDF text search API call. // See documentation: https://developer.pdf.co String query = "https://api.pdf.co/v1/pdf/find"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\", \"searchString\": \"%s\", \"regexSearch\": \"%s\"}", Password, Pages, SourceFileURL, SearchString, RegexSearch); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Display found items in console for (JsonElement element : json.get("body").getAsJsonArray()) { JsonObject item = (JsonObject) element; System.out.println("Found text " + item.get("text") + " at coordinates " + item.get("left") + ", "+ item.get("top")); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } } ``` # Find Text in Table with AI Source: https://developer.pdf.co/api/pdf-find/table Detect tables in PDF and scanned documents using AI, returning the page, coordinates, and detected column structure for each table. **Try it live:** [Find Text in Table with AI → API Tester](/api-tester/pdf-find/table) — send a real request from your browser. ## `POST /v1/pdf/find/table` This function finds tables in documents using an AI-powered table detection engine. This endpoint locates `tables` in an input PDF document and returns JSON with: * The array of tables objects. * `X`, `Y`, `Width`, and `Height` coordinates for every table found. * `Rect` param for every table that you can re-use with `pdf/convert/to/json`, `pdf/convert/to/csv`, `pdf/convert/to/csv`, and other endpoints to extract a selected table only. * `PageIndex` page index for a page with a table. The very first page is `0` . * `Columns` array with the set of `X` coordinates for every column inside the table that was found. To extract the table into CSV, JSON, or XML please use pdf/convert/to/csv, pdf/convert/to/json2, and pdf/convert/to/xml endpoints with rect parameter value from rect output param for this table accordingly. To extract the table into CSV, JSON, or XML please use [pdf/convert/to/csv](/api/pdf-to-csv), [pdf/convert/to/json2](/api/pdf-to-json/with-ai), and [pdf/convert/to/xml](/api/pdf-to-xml) endpoints with `rect` parameter value from `rect` output param for this table accordingly. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Find only bordered tables You can limit search to bordered tables only by enabling the *legacy table* search mode with the following `profiles` config: ```json theme={null} { "profiles": "{ 'Mode': 'Legacy', 'ColumnDetectionMode': 'BorderedTables', 'DetectionMinNumberOfRows': 1, 'DetectionMinNumberOfColumns': 1, 'DetectionMaxNumberOfInvalidSubsequentRowsAllowed': 0, 'DetectionMinNumberOfLineBreaksBetweenTables': 0, 'EnhanceTableBorders': false }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `body` | object | Response body. | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "async": "false", "inline": "true", "password": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "tables": [ { "PageIndex": 0, "X": 36, "Y": 34.4400024, "Width": 523.44, "Height": 160.82, "Columns": [ 357.675 ], "rect": "36, 34.4400024, 523.44, 160.82" }, { "PageIndex": 0, "X": 36, "Y": 316.249969, "Width": 523.44, "Height": 120.620026, "Columns": [ 157.117, 340.68, 475.84 ], "rect": "36, 316.249969, 523.44, 120.620026" } ] }, "pageCount": 1, "error": false, "status": 200, "name": "sample.json", "remainingCredits": 98892697, "credits": 21 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/find/table' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf", "async": "false", "inline": "true", "password": "" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Prepare URL for PDF Table Search API call. // See documentation: https://developer.pdf.co var query = `https://api.pdf.co/v1/pdf/find/table`; let reqOptions = { uri: query, headers: { "x-api-key": API_KEY }, formData: { password: Password, pages: Pages, url: SourceFileUrl } }; // Send request request.post(reqOptions, function (error, resp, body) { if (error) { return console.error("Error: ", error); } var jsonBody = JSON.parse(body); // Loop through all found tables, and get json data if (jsonBody.body.tables && jsonBody.body.tables.length > 0) { for (var i = 0; i < jsonBody.body.tables.length; i++) { getJSONFromCoordinates(SourceFileUrl, jsonBody.body.tables[i].PageIndex, jsonBody.body.tables[i].rect, `table_${i + 1}.json`); } } }); /** * Get JSON from specific co-ordinates */ function getJSONFromCoordinates(fileUrl, pageIndex, rect, outputFileName) { // Prepare request to `PDF To JSON` API endpoint var jsonQueryPath = `https://api.pdf.co/v1/pdf/convert/to/json`; // Json Request let jsonReqOptions = { uri: jsonQueryPath, headers: { "x-api-key": API_KEY }, formData: { pages: pageIndex, url: fileUrl, rect: rect } }; // Send request request.post(jsonReqOptions, function (error, resp, body) { if (error) { return console.error("Error: ", error); } var outputJsonUrl = JSON.parse(body).url; // Download JSON file var file = fs.createWriteStream(outputFileName); https.get(outputJsonUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated JSON file saved as "${outputFileName}" file.`); }); }); }); } ``` ```python theme={null} import requests import os # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***************************************" # Direct URL of source PDF file. SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Prepare URL for PDF Table Search API call. query = "https://api.pdf.co/v1/pdf/find/table" reqOptions = { 'password': Password, 'pages': Pages, 'url': SourceFileUrl } headers = { 'x-api-key': API_KEY } def getJSONFromCoordinates(fileUrl, pageIndex, rect, outputFileName): # Prepare request to `PDF To JSON` API endpoint jsonQueryPath = "https://api.pdf.co/v1/pdf/convert/to/json" # Json Request jsonReqOptions = { 'pages': pageIndex, 'url': fileUrl, 'rect': rect } # Send request response = requests.post(jsonQueryPath, headers=headers, data=jsonReqOptions) if response.status_code == 200: outputJsonUrl = response.json()['url'] # Download JSON file res = requests.get(outputJsonUrl) with open(outputFileName, 'wb') as outfile: outfile.write(res.content) print(f'Generated JSON file saved as "{outputFileName}" file.') else: print(f"Request error: {response.status_code} {response.reason}") # Send request response = requests.post(query, headers=headers, data=reqOptions) if response.status_code == 200: jsonBody = response.json() # Loop through all found tables, and get json data if 'tables' in jsonBody['body'] and len(jsonBody['body']['tables']) > 0: for i, table in enumerate(jsonBody['body']['tables']): getJSONFromCoordinates(SourceFileUrl, table['PageIndex'], table['rect'], f"table_{i + 1}.json") else: print(f"Request error: {response.status_code} {response.reason}") ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.Net; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*****************************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for PDF Table Search API call. // See documentation: https://developer.pdf.co string url = "https://api.pdf.co/v1/pdf/find/table"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", SourceFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["status"].ToString() != "error") { Console.WriteLine(response); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for PDF Table Search API call. // See documentation: https://developer.pdf.co String query = "https://api.pdf.co/v1/pdf/find/table"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", Password, Pages, SourceFileURL); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { System.out.println(response.body().string()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } } ``` # PDF from CSV Source: https://developer.pdf.co/api/pdf-from-document/csv Convert `CSV`, `XLS`, `XLSX` files into `PDF`. **Try it live:** [PDF from CSV → API Tester](/api-tester/pdf-from-document/csv) — send a real request from your browser. ## `POST /pdf/convert/from/csv` During conversion you should not expect any Word macros to operate as we do not support Office macros. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `autosize` | boolean | *No* | - | Set to `true` to page dimensions adjust to content with automatic page sizing. If false, uses worksheet's page setup. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "pages": "0-", "name": "result.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/efc283805b4a47da87910826d4ddf063/result.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEKz%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDFzXkfTapcUbLLKahiKBAbIL4F2wV3gvozuGDxmOpWUu9ETuVzkYKjMuNLAFzVZeSgRm9Yuaj7ubad9uOLQkL65GNgBQoy1Xm%2FxtLWD9tegUYd3hFvYfIWMfkWjuROwMGTZeD3CMacDPdFkP%2BUSG4aXOZb8MoG2PXnsd9UUeOvrevZkCVTg77OBXIteBCPOojSjeis%2F5BTIoVCi%2FrwV5kEGkbfBwtgsfQL3MxSbg7j%2Fud%2F3oGbUWW7zsemcfiHTiFg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAK44PZ6O/20200812/us-west-2/s3/aws4_request&X-Amz-Date=20200812T103301Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=6a176281828de74c917a4ff5bacd46eeca50221178bf34d01cac331f172e51c3", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 61165 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/doc' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "pages": "0-", "name": "result.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `DOC to PDF` API endpoint var queryPath = `/v1/pdf/convert/from/doc`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source DOC file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): convertDOCToPDF(SourceFileURL, DestinationFile) def convertDOCToPDF(uploadedFileUrl, destinationFile): """Converts DOC to PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrl # Prepare URL for 'DOC To PDF' API request url = "{}/pdf/convert/from/doc".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", SourceFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // URL of `DOC To PDF` API call string url = "https://api.pdf.co/v1/pdf/convert/from/doc"; try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `DOC To PDF` API call String query = "https://api.pdf.co/v1/pdf/convert/from/doc"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "DOC To PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF from DOC Source: https://developer.pdf.co/api/pdf-from-document/doc Convert `DOC`, `DOCX`, `RTF`, `TXT`, `XPS`, `PPT`, `PPTX` files into `PDF`. **Try it live:** [PDF from DOC → API Tester](/api-tester/pdf-from-document/doc) — send a real request from your browser. ## `POST /pdf/convert/from/doc` During conversion you should not expect any Word macros to operate as we do not support Office macros. This endpoint can be utilized as-is to convert PowerPoint files (`PPT`, `PPTX`) to PDF. However, `PPT` and `PPTX` **are not natively supported**, and we may not provide assistance or fixes for any issues, bugs, or limitations encountered during its use. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `autosize` | boolean | *No* | - | Set to `true` to page dimensions adjust to content with automatic page sizing. If false, uses worksheet's page setup. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "pages": "0-", "name": "result.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/efc283805b4a47da87910826d4ddf063/result.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEKz%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDFzXkfTapcUbLLKahiKBAbIL4F2wV3gvozuGDxmOpWUu9ETuVzkYKjMuNLAFzVZeSgRm9Yuaj7ubad9uOLQkL65GNgBQoy1Xm%2FxtLWD9tegUYd3hFvYfIWMfkWjuROwMGTZeD3CMacDPdFkP%2BUSG4aXOZb8MoG2PXnsd9UUeOvrevZkCVTg77OBXIteBCPOojSjeis%2F5BTIoVCi%2FrwV5kEGkbfBwtgsfQL3MxSbg7j%2Fud%2F3oGbUWW7zsemcfiHTiFg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAK44PZ6O/20200812/us-west-2/s3/aws4_request&X-Amz-Date=20200812T103301Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=6a176281828de74c917a4ff5bacd46eeca50221178bf34d01cac331f172e51c3", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 61165 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/doc' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx", "pages": "0-", "name": "result.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `DOC to PDF` API endpoint var queryPath = `/v1/pdf/convert/from/doc`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source DOC file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): convertDOCToPDF(SourceFileURL, DestinationFile) def convertDOCToPDF(uploadedFileUrl, destinationFile): """Converts DOC to PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrl # Prepare URL for 'DOC To PDF' API request url = "{}/pdf/convert/from/doc".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", SourceFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // URL of `DOC To PDF` API call string url = "https://api.pdf.co/v1/pdf/convert/from/doc"; try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source DOC or DOCX file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `DOC To PDF` API call String query = "https://api.pdf.co/v1/pdf/convert/from/doc"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "DOC To PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF from Email Source: https://developer.pdf.co/api/pdf-from-email Convert email files (`.msg` or `.eml`) code into PDF. Extract attachments (if any) from input email and embeds into PDF as PDF attachments. **Try it live:** [PDF from Email → API Tester](/api-tester/pdf-from-email) — send a real request from your browser. ## `POST /v1/pdf/convert/from/email` Images and attachments within `.eml` and `.msg` files **must be publicly accessible.** Resources stored on local file systems or gated behind authentication are not supported for processing. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | --------------------------------------------------------- | ------- | -------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | | `paperSize` | string | *No* | `A4` | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. | | `orientation` | string | *No* | `Portrait` | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `embedAttachments` | boolean | *No* | `true` | Set to true to automatically embeds all attachments from original input email MSG or EML files into the final output PDF. Set it to false if you don't want to embed attachments so it will convert only the body of the input email. True by default. | | `convertAttachments` | boolean | *No* | `true` | Set to false if you don't want to convert attachments from the original email and want to embed them as original files (as embedded PDF attachments). Converts attachments that are supported into PDF format and then merges into output final PDF. The supported attachment types for conversion are: `.eml`, `.html/.htm`, `.pdf`, `.doc`, `.docx`, and `.rtf`. Non-supported file types are added as PDF attachments (Adobe Reader or another viewer may be required to view PDF attachments). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     [`removeHTMLHeadStyleTags`](#removehtmlheadstyletags) | string | *No* | `false` | Removes default styles from the `` section of Outlook emails to ensure accurate PDF rendering. Set to `'true'` to enable. | |     [`removeHTMLBodyStyleTags`](#removehtmlbodystyletags) | string | *No* | `false` | Removes inline and default styles from the `` section of Outlook emails to ensure accurate PDF rendering. Set to `'true'` to enable. | |     [`CustomScript`](#customscript) | string | *No* | - | Custom JavaScript code executed on the email content before PDF conversion. Use to modify HTML elements, disable links, or apply custom transformations. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Profiles examples #### removeHTMLHeadStyleTags * **Remove default Outlook styles for accurate PDF conversion** ```json theme={null} { "profiles": { "removeHTMLHeadStyleTags": "true", } } ``` #### removeHTMLBodyStyleTags * **Remove inline and default styles for accurate PDF conversion** ```json theme={null} { "profiles": { "removeHTMLBodyStyleTags": "true" } } ``` This example removes all inline and default styles from the `` section of Outlook emails to ensure accurate PDF rendering. #### CustomScript ```json theme={null} { "profiles": { "CustomScript": "document.querySelectorAll('a').forEach(a => { a.href = '#' });" } } ``` This example disables all active links while keeping the link text visible. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml", "embedAttachments": true, "convertAttachments": true, "paperSize": "Letter", "name": "email-with-attachments", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/980bc13f061344809c75e83ce181851c/Contact_us.pdf", "pageCount": 3, "error": false, "status": 200, "name": "Contact_us.pdf", "remainingCredits": 60637 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/email' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml", "embedAttachments": true, "convertAttachments": true, "paperSize": "Letter", "name": "email-with-attachments", "async": false }' ``` ```javascript theme={null} var request = require('request'); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ var options = { 'method': 'POST', 'url': 'https://api.pdf.co/v1/pdf/convert/from/email', 'headers': { 'Content-Type': 'application/json', 'x-api-key': '{{x-api-key}}' }, formData: { 'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml', 'embedAttachments': 'true', 'convertAttachments': 'true', 'paperSize': 'Letter', 'name': 'email-with-attachments', 'async': 'false' } }; request(options, function (error, response) { if (error) throw new Error(error); console.log(response.body); }); ``` ```python theme={null} import requests url = "https://api.pdf.co/v1/pdf/convert/from/email" # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ payload={'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml', 'embedAttachments': 'true', 'convertAttachments': 'true', 'paperSize': 'Letter', 'name': 'email-with-attachments', 'async': 'false'} files=[ ] headers = { 'Content-Type': 'application/json', 'x-api-key': '{{x-api-key}}' } response = requests.request("POST", url, headers=headers, json=payload, files=files) print(response.text) ``` ```csharp theme={null} using Newtonsoft.Json; using Newtonsoft.Json.Linq; using System; using System.Collections.Generic; using System.Net; using System.Threading; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source Email file to convert // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = @"https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml"; // Ouput file path const string DestinationFile = @"output.pdf"; // (!) Make asynchronous job const bool Async = true; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); try { // URL for `PDF FROM Email` API call var url = "https://api.pdf.co/v1/pdf/convert/from/email"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); // Link to input EML or MSG file to be converted. // You can pass link to file from Google Drive, Dropbox or another online file service that can generate shareable links. // You can also use built-in PDF.co cloud storage located at https://app.pdf.co/files or upload your file as temporary file right before making this API call (see Upload and Manage Files section for more details on uploading files via API). parameters.Add("url", SourceFileUrl); // True by default. // Set to true to automatically embeds all attachments from original input email MSG or EML fileas files into final output PDF. // Set to false if you don’t want to embed attachments so it will convert only the body of input email. parameters.Add("embedAttachments", true); // true by default. // Converts attachments that are supported by API (doc, docx, html, png, jpg etc) into PDF and merges into output final PDF. // Non-supported file types are added as PDF attachments (Adobe Reader or another viewer maybe required to view PDF attachments). // Set to false if you don’t want to convert attachments from original email and want to embed them as original files (as embedded pdf attachments). parameters.Add("convertAttachments", true); // Can be Letter, A4, A5, A6 or custom size like 200x200 parameters.Add("paperSize", "Letter"); // Name of output PDF parameters.Add("name", "email-with-attachments"); // Set to true to run as async job in background (recommended for heavy documents). parameters.Add("async", Async); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Asynchronous job ID string jobId = json["jobId"].ToString(); // URL of generated JSON file available after the job completion; it will contain URLs of result PDF files. string resultFileUrl = json["url"].ToString(); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status); if (status == "success") { // Download output file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); break; } else if (status == "working") { // Pause for a few seconds Thread.Sleep(3000); } else { Console.WriteLine(status); break; } } while (true); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } /// /// Checks Job Status /// static string CheckJobStatus(string jobId) { using (WebClient webClient = new WebClient()) { // Set API Key webClient.Headers.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; string response = webClient.DownloadString(url); JObject json = JObject.Parse(response); return Convert.ToString(json["status"]); } } } } ``` ```java theme={null} import java.io.*; import okhttp3.*; public class main { public static void main(String []args) throws IOException{ OkHttpClient client = new OkHttpClient().newBuilder() .build(); MediaType mediaType = MediaType.parse("application/json"); // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM) .addFormDataPart("url","https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml") .addFormDataPart("embedAttachments","true") .addFormDataPart("convertAttachments","true") .addFormDataPart("paperSize","Letter") .addFormDataPart("name","email-with-attachments") .addFormDataPart("async","false") .build(); Request request = new Request.Builder() .url("https://api.pdf.co/v1/pdf/convert/from/email") .method("POST", body) .addHeader("Content-Type", "application/json") .addHeader("x-api-key", "{{x-api-key}}") .build(); Response response = client.newCall(request).execute(); System.out.println(response.body().string()); } } ``` ```php theme={null} 'https://api.pdf.co/v1/pdf/convert/from/email', CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => '', CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 0, CURLOPT_FOLLOWLOCATION => true, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => 'POST', CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml','embedAttachments' => 'true','convertAttachments' => 'true','paperSize' => 'Letter','name' => 'email-with-attachments','async' => 'false'), CURLOPT_HTTPHEADER => array( 'Content-Type: application/json', 'x-api-key: {{x-api-key}}' ), )); $response = json_decode(curl_exec($curl)); curl_close($curl); echo "

Output:

", var_export($response, true), "
"; ```
# PDF from HTML Source: https://developer.pdf.co/api/pdf-from-html/convert Convert raw HTML markup into a PDF document with custom margins, paper size, orientation, headers and footers. **Try it live:** [PDF from HTML → API Tester](/api-tester/pdf-from-html/convert) — send a real request from your browser. ## `POST /v1/pdf/convert/from/html` This API converts a RAW HTML code into a PDF document and process any JavaScript which the webpage triggers when it loads. For example if the the webpage triggers a JavaScript popup window then that will be included in the conversion process. There is no option to disable JavaScript on the supplied RAW HTML code. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` Remember to ensure that request sizes are less than `4` mb in file size. For more information, see [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `html` | string | *Yes* | - | Input HTML code to be converted. To convert the link to a PDF use the /pdf/convert/from/url endpoint instead. | | `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | | `paperSize` | string | *No* | A4 | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. | | `orientation` | string | *No* | Portrait | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | | `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL | | `mediaType` | string | *No* | print | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). | | `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. | | `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. | | `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Header & Footer The `header` and `footer` parameters can contain valid HTML markup with the following classes used to inject printing values into them: * `date`: formatted print date * `title`: document title * `url`: document location * `pageNumber`: current page number * `totalPages`: total pages in the document img: tag is supported in both the header and footer parameter, provided that the `src` attribute is specified as a `base64-encoded` string. For example, the following markup will generate `Page N of NN` page numbering: ```html theme={null} Page of . ``` ### Sample Header & Footer An example with an advanced `header` and `footer`. Note that the top and bottom page margins are important because page content may overlap the footer or header. ```json theme={null} { "html": "

Hello

", "async": false, "name": "result.pdf", "margins": "40px 5px 40px 5px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "
LEFT SUBHEADERRIGHT SUBHEADER
", "footer": "
Page of .
" } ``` If you use `JSON` as input then make sure to escape it first (with `JSON.stringify(dataObject)` in JS). Escaping is when every `"` is replaced with `\"`. Example with `"` be escaped as `\"` then: `"templateData": "{ \"paid\": true, \"invoice_id\": \"0002\", \"total\": \"$999.99\" }"`. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "html": "

Hello World!

Go to PDF.co", "name": "result.pdf", "margins": "5px 5px 5px 5px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "mediaType": "print", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 60646 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/html' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "html": "

Hello World!

Go to PDF.co", "name": "result.pdf", "margins": "5px 5px 5px 5px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "mediaType": "print", "async": false }' ```
```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***************************"; // HTML Input const inputHtml = "./sample.html"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare requests params as JSON var parameters = {}; // Input HTML code to be converted. Required. parameters["html"] = fs.readFileSync(inputHtml, "utf8"); // Name of resulting file parameters["name"] = path.basename(DestinationFile); // Set to css style margins like 10 px or 5px 5px 5px 5px. parameters["margins"] = "5px 5px 5px 5px"; // Can be Letter, A4, A5, A6 or custom size like 200x200 parameters["paperSize"] = "Letter"; // Set to Portrait or Landscape. Portrait by default. parameters["orientation"] = "Portrait"; // true by default. Set to false to disbale printing of background. parameters["printBackground"] = true; // If large input document, process in async mode by passing true parameters["async"] = false; // Set to HTML for header to be applied on every page at the header. parameters["header"] = ""; // Set to HTML for footer to be applied on every page at the bottom. parameters["footer"] = ""; // Convert JSON object to string var jsonPayload = JSON.stringify(parameters); // Prepare request to `HTML To PDF` API endpoint var url = '/v1/pdf/convert/from/html'; var reqOptions = { host: "api.pdf.co", path: url, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import json import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "**************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # HTML template file_read = open(".\\sample.html", mode='r', encoding= 'utf-8') SampleHtml = file_read.read() file_read.close() # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): GeneratePDFFromHtml(SampleHtml, DestinationFile) def GeneratePDFFromHtml(SampleHtml, destinationFile): """Converts HTML to PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} # Input HTML code to be converted. Required. parameters["html"] = SampleHtml # Name of resulting file parameters["name"] = os.path.basename(destinationFile) # Set to css style margins like 10 px or 5px 5px 5px 5px. parameters["margins"] = "5px 5px 5px 5px" # Can be Letter, A4, A5, A6 or custom size like 200x200 parameters["paperSize"] = "Letter" # Set to Portrait or Landscape. Portrait by default. parameters["orientation"] = "Portrait" # true by default. Set to false to disable printing of background. parameters["printBackground"] = "true" # If large input document, process in async mode by passing true parameters["async"] = "false" # Set to HTML for header to be applied on every page at the header. parameters["header"] = "" # Set to HTML for footer to be applied on every page at the bottom. parameters["footer"] = "" # Prepare URL for 'HTML To PDF' API request url = "{}/pdf/convert/from/html".format( BASE_URL ) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "******************************"; static void Main(string[] args) { // HTML input string inputSample = File.ReadAllText(@".\sample.html"); // Destination PDF file name string destinationFile = @".\result.pdf"; // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Set JSON content type webClient.Headers.Add("Content-Type", "application/json"); try { // Prepare requests params as JSON Dictionary parameters = new Dictionary(); // Input HTML code to be converted. Required. parameters.Add("html", inputSample); // Name of resulting file parameters.Add("name", Path.GetFileName(destinationFile)); // Set to css style margins like 10 px or 5px 5px 5px 5px. parameters.Add("margins", "5px 5px 5px 5px"); // Can be Letter, A4, A5, A6 or custom size like 200x200 parameters.Add("paperSize", "Letter"); // Set to Portrait or Landscape. Portrait by default. parameters.Add("orientation", "Portrait"); // true by default. Set to false to disbale printing of background. parameters.Add("printBackground", true); // If large input document, process in async mode by passing true parameters.Add("async", false); // Set to HTML for header to be applied on every page at the header. parameters.Add("header", ""); // Set to HTML for footer to be applied on every page at the bottom. parameters.Add("footer", ""); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Prepare URL for `HTML to PDF` API call string url = "https://api.pdf.co/v1/pdf/convert/from/html"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); webClient.Headers.Remove("Content-Type"); // remove the header required for only the previous request // Download the PDF file webClient.DownloadFile(resultFileUrl, destinationFile); Console.WriteLine("Generated PDF document saved as \"{0}\" file.", destinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key to exit..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { // HTML input final String inputSample = new String(Files.readAllBytes(Paths.get(".\\sample.html"))); // Destination PDF file name final Path destinationFile = Paths.get(".\\result.pdf"); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `HTML to PDF` API call String apiUrl = "https://api.pdf.co/v1/pdf/convert/from/html"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, apiUrl, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Prepare request body in JSON format JsonObject jsonBody = new JsonObject(); // Input HTML code to be converted. Required. jsonBody.add("html", new JsonPrimitive(inputSample)); // Name of resulting file jsonBody.add("name", new JsonPrimitive(destinationFile.getFileName().toString())); // Set to css style margins like 10 px or 5px 5px 5px 5px. jsonBody.add("margins", new JsonPrimitive("5px 5px 5px 5px")); // Can be Letter, A4, A5, A6 or custom size like 200x200 jsonBody.add("paperSize", new JsonPrimitive("Letter")); // Set to Portrait or Landscape. Portrait by default. jsonBody.add("orientation", new JsonPrimitive("Portrait")); // true by default. Set to false to disable printing of background. jsonBody.add("printBackground", new JsonPrimitive(true)); // If large input document, process in async mode by passing true jsonBody.add("async", new JsonPrimitive(false)); // Set to HTML for header to be applied on every page at the header. jsonBody.add("header", new JsonPrimitive("")); // Set to HTML for footer to be applied on every page at the bottom. jsonBody.add("footer", new JsonPrimitive("")); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} HTML to PDF Result Error: " . $json["message"] . "

"; } else { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***************************"; // HTML Input const inputHtml = "./sample.html"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare requests params as JSON var parameters = {}; // Input HTML code to be converted. Required. parameters["html"] = fs.readFileSync(inputHtml, "utf8"); // Name of resulting file parameters["name"] = path.basename(DestinationFile); // Set to css style margins like 10 px or 5px 5px 5px 5px. parameters["margins"] = "5px 5px 5px 5px"; // Can be Letter, A4, A5, A6 or custom size like 200x200 parameters["paperSize"] = "Letter"; // Set to Portrait or Landscape. Portrait by default. parameters["orientation"] = "Portrait"; // true by default. Set to false to disbale printing of background. parameters["printBackground"] = true; // If large input document, process in async mode by passing true parameters["async"] = false; // Set to HTML for header to be applied on every page at the header. parameters["header"] = ""; // Set to HTML for footer to be applied on every page at the bottom. parameters["footer"] = ""; // Convert JSON object to string var jsonPayload = JSON.stringify(parameters); // Prepare request to `HTML To PDF` API endpoint var url = '/v1/pdf/convert/from/html'; var reqOptions = { host: "api.pdf.co", path: url, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ```
# PDF from HTML Template Source: https://developer.pdf.co/api/pdf-from-html/convert-from-template Convert `HTML` template into `PDF`. ## `POST /v1/pdf/convert/from/html` Converts a predefined [HTML template](#html-templates) into a PDF document using its Template ID. The template must be created and saved in the [HTML to PDF Templates](https://app.pdf.co/html-templates-tool/manager) section of the dashboard. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` Remember to ensure that request sizes are less than `4` mb in file size. For more information, see [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `templateId` | integer | *Yes* | - | Set ID of HTML template to be used. View and manage your templates at HTML to PDF Templates. | | `templateData` | string | *Yes* | - | Set it to a string with input `JSON` data (recommended) or `CSV` data. See [Sample JSON input](#sample-json-input) and [Sample CSV input](#sample-csv-input) for more information. | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | | `paperSize` | string | *No* | A4 | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. | | `orientation` | string | *No* | Portrait | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | | `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL | | `mediaType` | string | *No* | print | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). | | `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. | | `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. | | `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## HTML Templates During conversion, the API performs variable substitution and built-in helper evaluation using your [templateData](#html-templates), but does not execute any arbitrary JavaScript within the template itself. After the template is rendered into HTML, a headless browser processes the page and executes any in-page JavaScript triggered on load. External or dynamically loaded JavaScript libraries may run if accessible, but execution is not guaranteed in all environments. This separation ensures secure, predictable template handling and accurate rendering of interactive web content. Use the dashboard to manage your HTML to [PDF Templates](https://app.pdf.co/html-templates-tool). Templates use `{{Mustache}}` and Handlebars templating syntax. You just need to insert macros surrounded by double brackets like `{{` and `}}`. * Find out more about [Mustache](https://mustache.github.io/mustache.5.html). * Find out more about [Handlebars](https://handlebarsjs.com/guide/). Some Examples of macro inside html template: * `{{variable1}}` will be replaced with `test` if you set `templateData` to `{ "variable1": "test" }` * `{{object1.variable1}}` will be replaced with `test` if you set `templateData` to `{ "object1": { "variable1": "test" } }` * Simple conditions are also supported. For example: `{{#if paid}} invoice was paid {{/if}}` will show invoice was paid when `templateData` is set to `{ "paid": true }`. ### Handlebars Handlebars extends Mustache with powerful features like conditional logic, loops, and custom helper functions. You can use Handlebars helpers like `#if`, `#unless`, `#each` (see [https://handlebarsjs.com/guide/builtin-helpers.html](https://handlebarsjs.com/guide/builtin-helpers.html)) but you can also define your own helper functions for complex calculations and data manipulation. #### Key Handlebars Features: * **Variable substitution**: `{{variable}}` syntax for inserting data * **Conditional logic**: `{{#if}}` and `{{#unless}}` blocks for conditional rendering * **Loops**: `{{#each}}` for iterating over arrays and objects * **Nested objects**: Accessing properties with dot notation like `{{company.name}}` * **Built-in helpers**: Using Handlebars' built-in functionality for common operations ### Sample JSON input ```json theme={null} "templateData": "{ 'paid': true, 'invoice_id': '0002', 'total': '$999.99' }" ``` If you use `JSON` as input then make sure to escape it first (with `JSON.stringify(dataObject)` in JS). Escaping is when every `"` is replaced with `\"`. Example with `"` be escaped as `\"` then: `"templateData": "{ \"paid\": true, \"invoice_id\": \"0002\", \"total\": \"$999.99\" }"`. ### Sample CSV input ```csv theme={null} "templateData": "paid,invoice_id,total true,0002,$999.99" ``` ## Header & Footer The `header` and `footer` parameters can contain valid HTML markup with the following classes used to inject printing values into them: * `date`: formatted print date * `title`: document title * `url`: document location * `pageNumber`: current page number * `totalPages`: total pages in the document img: tag is supported in both the header and footer parameter, provided that the `src` attribute is specified as a `base64-encoded` string. For example, the following markup will generate `Page N of NN` page numbering: ```html theme={null} Page of . ``` ### Sample Header & Footer An example with an advanced `header` and `footer`. Note that the top and bottom page margins are important because page content may overlap the footer or header. ```json theme={null} { "templateId": 1, "name": "newDocument.pdf", "mediaType": "print", "margins": "40px 20px 20px 20px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "
LEFT SUBHEADERRIGHT SUBHEADER
", "footer": "
Page of .
", "async": false, "templateData": "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "templateId": 1, "name": "newDocument.pdf", "mediaType": "print", "margins": "40px 20px 20px 20px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "async": false, "templateData": "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 60646 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/html' \ --header 'x-api-key: add_your_api_key_here' \ --header 'Content-Type: application/json' \ --data-raw '{ "templateId": 2, "name": "newDocument.pdf", "margins": "40px 20px 20px 20px", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "async": false, "encrypt": false, "templateData": "{\"invoice_id\":\"1234567\",\"invoice_date\":\"April 30, 2016\",\"invoice_dateDue\":\"May 15, 2016\",\"paid\":false,\"issuer_name\":\"Acme Inc\",\"issuer_company\":\"Acme International\",\"issuer_address\":\"City, Street 3rd\",\"issuer_email\":\"support@example.com\",\"issuer_website\":\"http://example.com\",\"client_name\":\"Food Delivery Inc.\",\"client_company\":\"Food Delivery International\",\"client_address\":\"New York, Some Street, 42\",\"client_email\":\"client@example.com\",\"items\":[{\"name\":\"Setting up new web-site\",\"price\":250},{\"name\":\"Website Content Addition\",\"price\":700},{\"name\":\"Database Setup\",\"price\":200},{\"name\":\"Record Digitalization\",\"price\":1800},{\"name\":\"Cloud Storage\",\"price\":500},{\"name\":\"Short Messages\",\"price\":35},{\"name\":\"Search Engine Optimization\",\"price\":200},{\"name\":\"Priority Support\",\"price\":75},{\"name\":\"Configuring mail server and mailboxes\",\"price\":50}],\"tax\":0.065,\"discount\":0.01,\"note\":\"Thank You For Your Business!\"}" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Data to fill the template const templateData = "./invoice_data.json"; // Destination PDF file name const DestinationFile = "./result.pdf"; /* Please follow below steps to create your own HTML Template and get "templateId". 1. Add new html template in app.pdf.co/templates/html 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" 3. Save this new template 4. Copy it’s ID to clipboard 5. Now set ID of the template into “templateId” parameter */ // HTML template using built-in template // see https://app.pdf.co/templates/html/2/edit const template_id = 2; // Prepare request to `HTML To PDF` API endpoint var queryPath = `/v1/pdf/convert/from/html?name=${path.basename(DestinationFile)}&async=True`; var reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json" } }; var requestBody = JSON.stringify({ "templateId": template_id, "templateData": fs.readFileSync(templateData, "utf8"), "async": true }); // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(requestBody); postRequest.end(); function checkIfJobIsCompleted(jobId, resultFileUrl) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***********************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # --HTML Template ID-- # Please follow below steps to create your own HTML Template and get "templateId". # 1. Add new html template in app.pdf.co/templates/html # 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" # 3. Save this new template # 4. Copy it’s ID to clipboard # 5. Now set ID of the template into “templateId” parameter # HTML template using built-in template # see https://app.pdf.co/templates/html/2/edit template_id = 2 # Data to fill the template file_read = open(".\\invoice_data.json", mode='r') TemplateData = file_read.read() file_read.close() # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): GeneratePDFFromTemplate(template_id, TemplateData, DestinationFile) def GeneratePDFFromTemplate(template_id, templateData, destinationFile): """Converts HTML to PDF using PDF.co Web API""" data = { 'templateData': templateData, 'templateId': template_id } # Prepare URL for 'HTML To PDF' API request url = "{}/pdf/convert/from/html?name={}".format( BASE_URL, os.path.basename(destinationFile) ) # Execute request and get response as JSON response = requests.post(url, data=data, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net.Http; using System.Threading.Tasks; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoCodeSample { class Program { const String API_KEY = "***********************************"; const string DestinationFile = @".\newDocument.pdf"; const bool Async = true; // Enable asynchronous processing static async Task Main(string[] args) { using (HttpClient httpClient = new HttpClient()) { httpClient.DefaultRequestHeaders.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/pdf/convert/from/html"; // Prepare requests params as JSON Dictionary parameters = new Dictionary { { "templateId", 1 }, { "name", Path.GetFileName(DestinationFile) }, { "margins", "40px 20px 20px 20px" }, { "paperSize", "Letter" }, { "orientation", "Portrait" }, { "header", "" }, { "printBackground", true }, { "footer", "" }, { "async", Async }, // Enable asynchronous processing { "encrypt", false }, { "templateData", "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}" } }; // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Send POST request asynchronously var content = new StringContent(jsonPayload, System.Text.Encoding.UTF8, "application/json"); HttpResponseMessage response = await httpClient.PostAsync(url, content); // Read response asynchronously string responseBody = await response.Content.ReadAsStringAsync(); JObject json = JObject.Parse(responseBody); if (json["error"].ToObject() == false) { // Get Job ID for asynchronous processing string jobId = json["jobId"].ToString(); Console.WriteLine($"Job ID: {jobId}"); // Check job status in a loop bool isJobCompleted = false; while (!isJobCompleted) { string status = await CheckJobStatusAsync(httpClient, jobId); Console.WriteLine($"Job Status: {status}"); if (status == "success") { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file asynchronously byte[] fileBytes = await httpClient.GetByteArrayAsync(resultFileUrl); await File.WriteAllBytesAsync(DestinationFile, fileBytes); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); isJobCompleted = true; } else if (status == "working") { // Wait for a few seconds before checking again await Task.Delay(3000); } else { Console.WriteLine($"Job failed or aborted. Status: {status}"); isJobCompleted = true; } } } else { Console.WriteLine(json["message"].ToString()); } } catch (Exception e) { Console.WriteLine(e.ToString()); } } Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } static async Task CheckJobStatusAsync(HttpClient httpClient, string jobId) { string url = $"https://api.pdf.co/v1/job/check?jobid={jobId}"; // Send GET request to check job status HttpResponseMessage response = await httpClient.GetAsync(url); string responseBody = await response.Content.ReadAsStringAsync(); JObject json = JObject.Parse(responseBody); return json["status"].ToString(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { /* Please follow below steps to create your own HTML Template and get "templateId". 1. Add new html template in app.pdf.co/templates/html 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" 3. Save this new template 4. Copy it’s ID to clipboard 5. Now set ID of the template into “templateId” parameter */ // HTML template using built-in template // see https://app.pdf.co/templates/html/2/edit final String templateId = "2"; // Data to fill the template final String templateData = new String(Files.readAllBytes(Paths.get(".\\invoice_data.json"))); // Destination PDF file name final Path destinationFile = Paths.get(".\\result.pdf"); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `HTML to PDF` API call String query = String.format( "https://api.pdf.co/v1/pdf/convert/from/html?name=%s", destinationFile.getFileName()); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Prepare request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("templateId", new JsonPrimitive(templateId)); jsonBody.add("templateData", new JsonPrimitive(templateData)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF Invoice Generation Results

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> ``` # Return HTML Template by ID Source: https://developer.pdf.co/api/pdf-from-html/template-id Returns HTML template by template’s id. **Try it live:** [Return HTML Template by ID → API Tester](/api-tester/pdf-from-html/template-id) — send a real request from your browser. ## `GET /templates/html/:id` Once you have obtained a template then use the [PDF from HTML](/api/pdf-from-html) API with the required `templateId` & `templateData` parameters defined. Use the dashboard to manage your [HTML to PDF Templates](https://app.pdf.co/html-templates-tool). ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------ | | `id` | integer | Template ID. | | `type` | string | Template type. | | `title` | string | Template title. | | `description` | string | Template description. | | `test_json` | string | Template test JSON. | | `updated_at` | string | Template updated at. | | `body` | string | Template content. | | `remainingCredits` | integer | Number of credits remaining in the account | | `credits` | integer | Number of credits consumed by the request | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "templates": [ { "id": 1, "type": "system", "title": "General Invoice Template", "description": "sample invoice template showcasing use of Mustache templates syntax for generating invoices" }, { "id": 15, "type": "user", "title": "User Template 1", "description": "" } ], "remainingCredits": 99204004, "credits": 2 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/templates/html' \ --header 'Content-Type: application/json' \ --header 'x-api-key: ' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Data to fill the template const templateData = "./invoice_data.json"; // Destination PDF file name const DestinationFile = "./result.pdf"; /* Please follow below steps to create your own HTML Template and get "templateId". 1. Add new html template in app.pdf.co/templates/html 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" 3. Save this new template 4. Copy it’s ID to clipboard 5. Now set ID of the template into “templateId” parameter */ // HTML template using built-in template // see https://app.pdf.co/templates/html/2/edit const template_id = 2; // Prepare request to `HTML To PDF` API endpoint var queryPath = `/v1/pdf/convert/from/html?name=${path.basename(DestinationFile)}&async=True`; var reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json" } }; var requestBody = JSON.stringify({ "templateId": template_id, "templateData": fs.readFileSync(templateData, "utf8"), "async": true }); // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(requestBody); postRequest.end(); function checkIfJobIsCompleted(jobId, resultFileUrl) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***********************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # --HTML Template ID-- # Please follow below steps to create your own HTML Template and get "templateId". # 1. Add new html template in app.pdf.co/templates/html # 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" # 3. Save this new template # 4. Copy it’s ID to clipboard # 5. Now set ID of the template into “templateId” parameter # HTML template using built-in template # see https://app.pdf.co/templates/html/2/edit template_id = 2 # Data to fill the template file_read = open(".\\invoice_data.json", mode='r') TemplateData = file_read.read() file_read.close() # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): GeneratePDFFromTemplate(template_id, TemplateData, DestinationFile) def GeneratePDFFromTemplate(template_id, templateData, destinationFile): """Converts HTML to PDF using PDF.co Web API""" data = { 'templateData': templateData, 'templateId': template_id } # Prepare URL for 'HTML To PDF' API request url = "{}/pdf/convert/from/html?name={}".format( BASE_URL, os.path.basename(destinationFile) ) # Execute request and get response as JSON response = requests.post(url, data=data, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; static void Main(string[] args) { // --TemplateID-- /* Please follow below steps to create your own HTML Template and get "templateId". 1. Add new html template in app.pdf.co/templates/html 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" 3. Save this new template 4. Copy it’s ID to clipboard 5. Now set ID of the template into “templateId” parameter */ // HTML template using built-in template // see https://app.pdf.co/templates/html/2/edit var templateId = 2; // Data to fill the template string templateData = File.ReadAllText(@".\invoice_data.json"); // Destination PDF file name string destinationFile = @".\result.pdf"; // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); webClient.Headers.Add("Content-Type", "application/json"); try { // URL for `HTML to PDF` API call string url = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/pdf/convert/from/html?name={0}", Path.GetFileName(destinationFile))); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(destinationFile)); parameters.Add("templateId", templateId); parameters.Add("templateData", templateData); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute request string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); webClient.Headers.Remove("Content-Type"); // remove the header required for only the previous request // Download the PDF file webClient.DownloadFile(resultFileUrl, destinationFile); Console.WriteLine("Generated PDF document saved as \"{0}\" file.", destinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key to exit..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import com.google.gson.JsonPrimitive; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; public static void main(String[] args) throws IOException { /* Please follow below steps to create your own HTML Template and get "templateId". 1. Add new html template in app.pdf.co/templates/html 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES" 3. Save this new template 4. Copy it’s ID to clipboard 5. Now set ID of the template into “templateId” parameter */ // HTML template using built-in template // see https://app.pdf.co/templates/html/2/edit final String templateId = "2"; // Data to fill the template final String templateData = new String(Files.readAllBytes(Paths.get(".\\invoice_data.json"))); // Destination PDF file name final Path destinationFile = Paths.get(".\\result.pdf"); // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `HTML to PDF` API call String query = String.format( "https://api.pdf.co/v1/pdf/convert/from/html?name=%s", destinationFile.getFileName()); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Prepare request body in JSON format JsonObject jsonBody = new JsonObject(); jsonBody.add("templateId", new JsonPrimitive(templateId)); jsonBody.add("templateData", new JsonPrimitive(templateData)); RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF Invoice Generation Results

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> ``` # Return All Templates Source: https://developer.pdf.co/api/pdf-from-html/templates Create PDF’s from HTML template input. **Try it live:** [Return All Templates → API Tester](/api-tester/pdf-from-html/templates) — send a real request from your browser. ## `GET /templates/html` Once you have obtained a template then use the [PDF from HTML](/api/pdf-from-html) API with the required templateId & templateData parameters defined. Use the dashboard to manage your [HTML to PDF Templates](https://app.pdf.co/html-templates-tool). ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | -------------- | ------------------------------------------ | | `templates` | array\[object] | | | `remainingCredits` | integer | Number of credits remaining in the account | | `credits` | integer | Number of credits consumed by the request | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "id": 1, "type": "system", "title": "General Invoice Template", "description": "sample invoice template showcasing use of Mustache templates syntax for generating invoices", "body": "\r\n\r\n\r\nInvoice \r\n\r\n \r\n\r\n \r\n
\r\n PAID
\r\n \r\n \r\n
\r\n
\r\n
\r\n \r\n \r\n
\r\n
\r\n
\r\n\r\n
\r\n
\r\n
\r\n
\r\n
\r\n
\r\n
\r\n
\r\n Invoice Number: \r\n
\r\n
\r\n Invoice Date: \r\n
\r\n
\r\n Invoice Due Date: \r\n
\r\n
\r\n
\r\n
\r\n \r\n
\r\n
\r\n\r\n
\r\n
BILL TO
\r\n
\r\n
Name:
\r\n
Company:
\r\n
Address:
\r\n
Email:
\r\n
\r\n
\r\n
\r\n \r\n
\r\n
\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n
ItemPrice
\r\n
\r\n\r\n
\r\n
\r\n
\r\n
\r\n
\r\n
Discount:
\r\n
Tax:
\r\n
TOTAL:
\r\n
\r\n \r\n
\r\n
\r\n
\r\n
\r\n
\r\n
\r\n\r\n\r\n​", "remainingCredits": 99204002, "credits": 2 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request GET 'https://api.pdf.co/v1/templates/html/1' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '' ``` # PDF from Image Source: https://developer.pdf.co/api/pdf-from-image Convert `JPG`, `PNG`, `TIFF` image formats into `PDF`. **Try it live:** [PDF from Image → API Tester](/api-tester/pdf-from-image) — send a real request from your browser. ## `POST /pdf/convert/from/image` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | To add multiple images to convert pleaase comma-separate the `url` parameter. e.g. "url": `"https://example.com/image1.png,https://example.com/image2.png"` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png,https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/5ef3d4033e344ec091bbeb3cfd848633/image2.pdf", "pageCount": 2, "error": false, "status": 200, "name": "image2.pdf", "remainingCredits": 59871 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/image' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png,https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URLs of image files to convert to PDF document // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFiles = [ "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" ]; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Image To PDF` API endpoint var queryPath = `/v1/pdf/convert/from/image`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFiles.join(",") }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URLs of image files to convert to PDF document # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFiles = [ "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" ] # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): SourceFileURL = ",".join(SourceFiles) convertImageToPDF(SourceFileURL, DestinationFile) def convertImageToPDF(uploadedFileUrl, destinationFile): """Converts Image to PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrl # Prepare URL for 'Image To PDF' API request url = "{}/pdf/convert/from/image".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URLs of image files to convert to PDF document // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ static string[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" }; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Image To PDF` API call string url = "https://api.pdf.co/v1/pdf/convert/from/image"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", string.Join(",", SourceFiles)); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URLs of image files to convert to PDF document // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String[] SourceFiles = { "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" }; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Image To PDF` API call String query = "https://api.pdf.co/v1/pdf/convert/from/image"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), String.join(",", SourceFiles)); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Image To PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF from URL Source: https://developer.pdf.co/api/pdf-from-url Convert any web page URL into a PDF document with custom margins, paper size, orientation, headers and footers. **Try it live:** [PDF from URL → API Tester](/api-tester/pdf-from-url) — send a real request from your browser. ## `POST /v1/pdf/convert/from/url` This method will process any JavaScript which the webpage triggers when it loads. For example if the the webpage triggers a JavaScript popup window then that will be included in the conversion process. There is no option to disable JavaScript on the supplied `HTML` page. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | | `paperSize` | string | *No* | `A4` | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. | | `orientation` | string | *No* | `Portrait` | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | | `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL | | `mediaType` | string | *No* | `print` | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). | | `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. | | `renderTimeout` | integer | *No* | - | Specifies the maximum time (in milliseconds) to wait for the page to fully render before generating the PDF. Useful for pages with heavy JavaScript, dynamic content, or slow-loading resources. The maximum allowed value is 177000ms (177 seconds). | | `legacyRendition` | boolean | *No* | `false` | Forces rendering with an older Chromium version (`r970485`) to work around version-specific issues where modern browsers can't paginate certain page structures correctly. Ensures full content is rendered as a long, printable page. | | `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. | | `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Header & Footer The header and footer parameters can contain valid HTML markup with the following classes used to inject printing values into them: date: formatted print date title: document title url: document location pageNumber: current page number totalPages: total pages in the document img: tag is supported in both the header and footer parameter, provided that the src attribute is specified as a base64-encoded string. For example, the following markup will generate Page N of NN page numbering: ```html theme={null} Page of . ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://wikipedia.org/wiki/Wikipedia:Contact_us", "name": "result.pdf", "margins": "5mm", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "mediaType": "print", "renderTimeout": 15000, "async": false, "profiles": "{ \"CustomScript\": \";; // put some custom js script here \"}" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf", "pageCount": 1, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 60646 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/url' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://wikipedia.org/wiki/Wikipedia:Contact_us", "margins": "5mm", "paperSize": "Letter", "orientation": "Portrait", "printBackground": true, "header": "", "footer": "", "mediaType": "print", "async": false, "profiles": "{ \"CustomScript\": \";; // put some custom js script here \"}" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // URL of web page to convert to PDF document. const SourceUrl = "http://en.wikipedia.org/wiki/Main_Page"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Web Page to PDF` API endpoint var queryPath = `/v1/pdf/convert/from/url`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "**********************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # URL of web page to convert to PDF document. SourceUrl = "http://en.wikipedia.org/wiki/Main_Page" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): convertHTMLToPDF(SourceUrl, DestinationFile) def convertHTMLToPDF(uploadedFileUrl, destinationFile): """Converts HTML to PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["url"] = uploadedFileUrl # Prepare URL for 'HTML To PDF' API request url = "{}/pdf/convert/from/url".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // URL of web page to convert to PDF document. const string SourceUrl = "http://en.wikipedia.org/wiki/Main_Page"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Web Page to PDF` API call string url = "https://api.pdf.co/v1/pdf/convert/from/url"; // Prepare requests params as JSON Dictionary requestBody = new Dictionary(); requestBody.Add("name", Path.GetFileName(DestinationFile)); requestBody.Add("url", SourceUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(requestBody); try { // Execute POST request var response = webClient.UploadString(url, "POST", jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF document saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // URL of web page to convert to PDF document. final static String SourceUrl = "http://en.wikipedia.org/wiki/Main_Page"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Web Page to PDF` API call String query = "https://api.pdf.co/v1/pdf/convert/from/url"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), SourceUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF Extractor Results

Conversion Result:

" . $resultFileUrl . ""; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); ?> ```
# PDF Info Reader Source: https://developer.pdf.co/api/pdf-info-reader Get detailed information about a **PDF** document, it’s properties and security permissions. **Try it live:** [PDF Info Reader → API Tester](/api-tester/pdf-info-reader) — send a real request from your browser. ## `POST /v1/pdf/info` Extracts basic information about an input PDF file, PDF file security permissions, and other information. If you want to extract information about fillable fields (checkboxes, radiobuttons, listboxes) from PDF then please use [/pdf/info/fields](/api/forms/info-reader) instead. For one-time check of PDF file information and find form fields please use PDF [Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters.AddGammaCorrection()` | array\[string (float format)] | *No* | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. | |     `OCRImagePreprocessingFilters.AddGrayscale()` | boolean | *No* | `false` | Set to true to preprocessing filter that converts a colored document/image to grayscale before performing OCR | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `info` | object | Info details. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "info": { "PageCount": 1, "Author": "Alice V. Knox", "Title": "Kid's News 1", "Producer": "Acrobat Distiller 4.0 for Windows", "Subject": "Kid's News 1", "CreationDate": "8/15/2001 2:50:36 PM", "Bookmarks": "", "Keywords": "", "Creator": "Adobe PageMaker 6.52", "Encrypted": false, "PageRectangle": { "Location": { "IsEmpty": true, "X": 0, "Y": 0 }, "Size": "612, 792", "X": 0, "Y": 0, "Width": 612, "Height": 792, "Left": 0, "Top": 0, "Right": 612, "Bottom": 792, "IsEmpty": false }, "ModificationDate": "9/20/2001 6:23:02 PM", "EncryptionAlgorithm": 0, "PermissionPrinting": true, "PermissionModifyDocument": true, "PermissionContentExtraction": true, "PermissionModifyAnnotations": true, "PermissionFillForms": true, "PermissionAccessibility": true, "PermissionAssemble": true, "PermissionHighQualityPrint": true }, "error": false, "status": 200, "remainingCredits": 77732 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/info' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file to get information const SourceFile = "./sample.pdf"; // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. GET INFORMATION FROM UPLOADED FILE getPdfInfo(API_KEY, uploadedFileUrl); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function getPdfInfo(apiKey, uploadedFileUrl) { // Prepare URL for `PDF Info` API call var queryPath = `/v1/pdf/info`; // JSON payload for api request var jsonPayload = JSON.stringify({ url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Display PDF document information for (var key in data.info) { console.log(`${key}: ${data.info[key]}`); } } else { // Service reported error console.log("getPdfInfo(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPdfInfo(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): getInfoFromPDF(uploadedFileUrl) def getInfoFromPDF(uploadedFileUrl): """Get Information using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF Info' API request url = "{}/pdf/info".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Display information print(json["info"]) else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file to get information const string SourceFile = @".\sample.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. GET INFORMATION FROM UPLOADED FILE // URL for `PDF Info` API call var url = "https://api.pdf.co/v1/pdf/info"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Display PDF document information foreach (JToken token in json["info"]) { JProperty property = (JProperty) token; Console.WriteLine("{0}: {1}", property.Name, property.Value); } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; import java.util.Map; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source file name final static Path SourceFile = Paths.get(".\\sample.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, uploadUrl, SourceFile.toFile())) { // 3. GET INFORMATION FROM UPLOADED FILE getPdfInfo(webClient, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void getPdfInfo(OkHttpClient webClient, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF Info` API call String query = "https://api.pdf.co/v1/pdf/info"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"url\": \"%s\"}", uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Display PDF document information JsonObject info = (JsonObject) json.get("info"); for (Map.Entry entry : info.entrySet()) { System.out.println(entry.getKey() + ": " + entry.getValue()); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String url, File sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } } ``` ```php theme={null} PDF Information Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractInfo($apiKey, $uploadedFileUrl) { // Create URL $url = "https://api.pdf.co/v1/pdf/info"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $documentInfo = $json["info"]; // Display the document info echo "

Document Info:

"; foreach ($documentInfo as $key => $value) { if(is_array($value)){ echo $key . ' = ' . json_encode($value) . '
'; } else{ echo $key . ' = ' . $value . '
'; } } echo "

"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } ?> ```
# Add Password to PDF Source: https://developer.pdf.co/api/pdf-password/add Add password and security limitations to PDF **Try it live:** [Add Password to PDF → API Tester](/api-tester/pdf-password/add) — send a real request from your browser. ## `POST /v1/pdf/security/add` **Modifying Restriction Settings**: To modify assembly or extraction settings that control PDF restrictions (such as `allowPrintDocument`, `allowFillForms`, `allowModifyDocument`, `allowAssemblyDocument`, and related permission settings), you **must** use the `ownerPassword` in your request. The `userPassword` alone cannot modify these permission settings. Attempting to change these restrictions with only a `userPassword` will result in an error: `"This file is password-protected. Please ensure you've entered the correct password…"`. This requirement applies only when making changes to those restrictions. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `ownerPassword` | string | *No* | - | The main owner password that is used for document encryption and for setting/removing restrictions. | | `userPassword` | string | *No* | - | The optional user password will be asked for viewing and printing document. | | `encryptionAlgorithm` | string | *No* | AES\_128bit | Encryption algorithm. AES\_128bit or higher is recommended. The available algorithms are: `RC4_40bit`, `RC4_128bit`, `AES_128bit`, `AES_256bit`. | | `allowAccessibilitySupport` | boolean | *No* | `false` | Allow or prohibit content extraction for accessibility needs. | | `allowAssemblyDocument` | boolean | *No* | `false` | Allow or prohibit assembling the document. | | `allowPrintDocument` | boolean | *No* | `false` | Allow or prohibit printing PDF document. | | `allowFillForms` | boolean | *No* | `false` | Allow or prohibit the filling of interactive form fields (including signature fields) in the PDF documents. | | `allowModifyDocument` | boolean | *No* | `false` | Allow or prohibit modification of PDF document. | | `allowContentExtraction` | boolean | *No* | `false` | Allow or prohibit copying content from PDF document. | | `allowModifyAnnotations` | boolean | *No* | `false` | Allow or prohibit interacting with text annotations and forms in PDF document. | | `printQuality` | string | *No* | HighResolution | Allowed printing quality. The available modes are: `LowResolution`, `HighResolution`. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | This restriction applies when `userPassword` (if any) is entered. This restriction does not apply if the user enters `ownerPassword`. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "ownerPassword": "12345", "userPassword": "54321", "EncryptionAlgorithm": "AES_128bit", "AllowPrintDocument": false, "AllowFillForms": false, "AllowModifyDocument": false, "AllowContentExtraction": false, "AllowModifyAnnotations": false, "PrintQuality": "LowResolution", "name": "output-protected.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/eaa441ade38548b8a3a96d8014c4f463/sample1.pdf", "pageCount": 1, "error": false, "status": 200, "name": "sample1.pdf", "remainingCredits": 616208, "credits": 14 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/security/add' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf", "ownerPassword": "12345", "userPassword": "54321", "EncryptionAlgorithm": "AES_128bit", "AllowPrintDocument": false, "AllowFillForms": false, "AllowModifyDocument": false, "AllowContentExtraction": false, "AllowModifyAnnotations": false, "PrintQuality": "LowResolution", "name": "output-protected.pdf", "async": false }' ``` # Remove Password from PDF Source: https://developer.pdf.co/api/pdf-password/remove Remove existing limits and password from PDF file. **Try it live:** [Remove Password from PDF → API Tester](/api-tester/pdf-password/remove) — send a real request from your browser. ## `POST /v1/pdf/security/remove` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-security/ProtectedPDFFile.pdf", "password": "admin@123", "name": "unprotected", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/9f2a754f76db46ac93781b3d2c6694c3/ProtectedPDFFile.pdf", "pageCount": 1, "error": false, "status": 200, "name": "ProtectedPDFFile.pdf", "remainingCredits": 616187, "credits": 21 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/security/remove' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-security/ProtectedPDFFile.pdf", "password": "admin@123", "name": "unprotected", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf"; // Destination PDF file name const DestinationFile = "./protected.pdf"; // Passwords to protect PDF document // The owner password will be required for document modification. // The user password only allows to view and print the document. const OwnerPassword = "123456"; const UserPassword = "654321"; // Encryption algorithm. // Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit". const EncryptionAlgorithm = "AES_128bit"; // Allow or prohibit content extraction for accessibility needs. const AllowAccessibilitySupport = true; // Allow or prohibit assembling the document. const AllowAssemblyDocument = true; // Allow or prohibit printing PDF document. const AllowPrintDocument = true; // Allow or prohibit filling of interactive form fields (including signature fields) in PDF document. const AllowFillForms = true; // Allow or prohibit modification of PDF document. const AllowModifyDocument = true; // Allow or prohibit copying content from PDF document. const AllowContentExtraction = true; // Allow or prohibit interacting with text annotations and forms in PDF document. const AllowModifyAnnotations = true; // Allowed printing quality. // Valid values: "HighResolution", "LowResolution" const PrintQuality = "HighResolution"; // Runs processing asynchronously. // Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success). const async = false; // Prepare request to `PDF Security` API endpoint var queryPath = `/v1/pdf/security/add`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), url: SourceFileUrl, ownerPassword: OwnerPassword, userPassword: UserPassword, encryptionAlgorithm: EncryptionAlgorithm, allowAccessibilitySupport: AllowAccessibilitySupport, allowAssemblyDocument: AllowAssemblyDocument, allowPrintDocument: AllowPrintDocument, allowFillForms: AllowFillForms, allowModifyDocument: AllowModifyDocument, allowContentExtraction: AllowContentExtraction, allowModifyAnnotations: AllowModifyAnnotations, printQuality: PrintQuality, async: async }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "********************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf" # Destination PDF file name DestinationFile = ".\\protected.pdf" # Passwords to protect PDF document # The owner password will be required for document modification. # The user password only allows to view and print the document. OwnerPassword = "123456" UserPassword = "654321" # Encryption algorithm. # Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit". EncryptionAlgorithm = "AES_128bit" # Allow or prohibit content extraction for accessibility needs. AllowAccessibilitySupport = True # Allow or prohibit assembling the document. AllowAssemblyDocument = True # Allow or prohibit printing PDF document. AllowPrintDocument = True # Allow or prohibit filling of interactive form fields (including signature fields) in PDF document. AllowFillForms = True # Allow or prohibit modification of PDF document. AllowModifyDocument = True # Allow or prohibit copying content from PDF document. AllowContentExtraction = True # Allow or prohibit interacting with text annotations and forms in PDF document. AllowModifyAnnotations = True # Allowed printing quality. # Valid values: "HighResolution", "LowResolution" PrintQuality = "HighResolution" # Runs processing asynchronously. # Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success). Async = False def main(args = None): protectPDF(SourceFileURL, DestinationFile) def protectPDF(uploadedFileUrl, destinationFile): """Protect PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {"name": os.path.basename(destinationFile), "url": uploadedFileUrl, "ownerPassword": OwnerPassword, "userPassword": UserPassword, "encryptionAlgorithm": EncryptionAlgorithm, "allowAccessibilitySupport": AllowAccessibilitySupport, "allowAssemblyDocument": AllowAssemblyDocument, "allowPrintDocument": AllowPrintDocument, "allowFillForms": AllowFillForms, "allowModifyDocument": AllowModifyDocument, "allowContentExtraction": AllowContentExtraction, "allowModifyAnnotations": AllowModifyAnnotations, "printQuality": PrintQuality, "async": Async} # Serializing json import json json_object = json.dumps(parameters, indent=4) # Prepare URL for 'PDF Security' API request url = "{}/pdf/security/add".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=json_object, headers={"x-api-key": API_KEY}) if (response.status_code == 200): jsonResp = response.json() if jsonResp["error"] == False: # Get URL of result file resultFileUrl = jsonResp["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(jsonResp["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.CodeDom; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Destination PDF file name const string DestinationFile = @".\protected.pdf"; // Passwords to protect PDF document // The owner password will be required for document modification. // The user password only allows to view and print the document. const string OwnerPassword = "123456"; const string UserPassword = "654321"; // Encryption algorithm. // Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit". const string EncryptionAlgorithm = "AES_128bit"; // Allow or prohibit content extraction for accessibility needs. const bool AllowAccessibilitySupport = true; // Allow or prohibit assembling the document. const bool AllowAssemblyDocument = true; // Allow or prohibit printing PDF document. const bool AllowPrintDocument = true; // Allow or prohibit filling of interactive form fields (including signature fields) in PDF document. const bool AllowFillForms = true; // Allow or prohibit modification of PDF document. const bool AllowModifyDocument = true; // Allow or prohibit copying content from PDF document. const bool AllowContentExtraction = true; // Allow or prohibit interacting with text annotations and forms in PDF document. const bool AllowModifyAnnotations = true; // Allowed printing quality. // Valid values: "HighResolution", "LowResolution" const string PrintQuality = "HighResolution"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Upload file to the cloud string uploadedFileUrl = UploadFile(SourceFile); // PROTECT UPLOADED PDF DOCUMENT // Prepare requests params as JSON // See documentation: https://developer.pdf.co/ Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("url", uploadedFileUrl); parameters.Add("ownerPassword", OwnerPassword); parameters.Add("userPassword", UserPassword); parameters.Add("encryptionAlgorithm", EncryptionAlgorithm); parameters.Add("allowAccessibilitySupport", AllowAccessibilitySupport.ToString()); parameters.Add("allowAssemblyDocument", AllowAssemblyDocument.ToString()); parameters.Add("allowPrintDocument", AllowPrintDocument.ToString()); parameters.Add("allowFillForms", AllowFillForms.ToString()); parameters.Add("allowModifyDocument", AllowModifyDocument.ToString()); parameters.Add("allowContentExtraction", AllowContentExtraction.ToString()); parameters.Add("allowModifyAnnotations", AllowModifyAnnotations.ToString()); parameters.Add("printQuality", PrintQuality); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // URL of "PDF Security" endpoint string url = "https://api.pdf.co/v1/pdf/security/add"; // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download generated PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } /// /// Uploads file to the cloud and return URL of uploaded file to use in further API calls. /// /// Source file name (path). /// URL of uploaded file static string UploadFile(string file) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); try { // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(file))); // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); // Get URL of uploaded file to use with later API calls string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", file); // You can use UploadData() instead if your file is in byte[] or Stream return uploadedFileUrl; } else { // Display service reported error Console.WriteLine(json["message"].ToString()); } } catch (Exception e) { Console.WriteLine(e); throw; } finally { webClient.Dispose(); } return null; } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\protected.pdf"); // Passwords to protect PDF document // The owner password will be required for document modification. // The user password only allows to view and print the document. final static String OwnerPassword = "123456"; final static String UserPassword = "654321"; // Encryption algorithm. // Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit". final static String EncryptionAlgorithm = "AES_128bit"; // Allow or prohibit content extraction for accessibility needs. final static boolean AllowAccessibilitySupport = true; // Allow or prohibit assembling the document. final static boolean AllowAssemblyDocument = true; // Allow or prohibit printing PDF document. final static boolean AllowPrintDocument = true; // Allow or prohibit filling of interactive form fields (including signature fields) in PDF document. final static boolean AllowFillForms = true; // Allow or prohibit modification of PDF document. final static boolean AllowModifyDocument = true; // Allow or prohibit copying content from PDF document. final static boolean AllowContentExtraction = true; // Allow or prohibit interacting with text annotations and forms in PDF document. final static boolean AllowModifyAnnotations = true; // Allowed printing quality. // Valid values: "HighResolution", "LowResolution" final static String PrintQuality = "HighResolution"; // Runs processing asynchronously. // Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success). final static boolean async = false; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `PDF Security` API call String query = "https://api.pdf.co/v1/pdf/security/add"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\n" + " \"name\": \"%s\",\n" + " \"url\": \"%s\",\n" + " \"ownerPassword\": \"%s\",\n" + " \"userPassword\": \"%s\",\n" + " \"EncryptionAlgorithm\": \"%s\",\n" + " \"AllowAccessibilitySupport\": %s,\n" + " \"AllowAssemblyDocument\": %s,\n" + " \"AllowPrintDocument\": %s,\n" + " \"AllowFillForms\": %s,\n" + " \"AllowModifyDocument\": %s,\n" + " \"AllowContentExtraction\": %s,\n" + " \"AllowModifyAnnotations\": %s,\n" + " \"PrintQuality\": \"%s\",\n" + " \"async\": %s\n" + "}", DestinationFile.getFileName(), SourceFileUrl, OwnerPassword, UserPassword, EncryptionAlgorithm, Boolean.toString(AllowAccessibilitySupport), Boolean.toString(AllowAssemblyDocument), Boolean.toString(AllowPrintDocument), Boolean.toString(AllowFillForms), Boolean.toString(AllowModifyDocument), Boolean.toString(AllowContentExtraction), Boolean.toString(AllowModifyAnnotations),PrintQuality, Boolean.toString(async) ); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ProtectPdf($apiKey, $uploadedFileUrl) { // Passwords to protect PDF document // The owner password will be required for document modification. // The user password only allows to view and print the document. $OwnerPassword = "123456"; $UserPassword = "654321"; // Encryption algorithm. // Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit". $EncryptionAlgorithm = "AES_128bit"; // Allow or prohibit content extraction for accessibility needs. $AllowAccessibilitySupport = true; // Allow or prohibit assembling the document. $AllowAssemblyDocument = true; // Allow or prohibit printing PDF document. $AllowPrintDocument = true; // Allow or prohibit filling of interactive form fields (including signature fields) in PDF document. $AllowFillForms = true; // Allow or prohibit modification of PDF document. $AllowModifyDocument = true; // Allow or prohibit copying content from PDF document. $AllowContentExtraction = true; // Allow or prohibit interacting with text annotations and forms in PDF document. $AllowModifyAnnotations = true; // Allowed printing quality. // Valid values: "HighResolution", "LowResolution" $PrintQuality = "HighResolution"; // Prepare URL for `PDF Security` API call $url = "https://api.pdf.co/v1/pdf/security/add"; // Prepare requests params $parameters = array(); $parameters["name"] = "result.pdf"; $parameters["url"] = $uploadedFileUrl; $parameters["ownerPassword"] = $OwnerPassword; $parameters["userPassword"] = $UserPassword; $parameters["encryptionAlgorithm"] = $EncryptionAlgorithm; $parameters["allowAccessibilitySupport"] = $AllowAccessibilitySupport; $parameters["allowAssemblyDocument"] = $AllowAssemblyDocument; $parameters["allowPrintDocument"] = $AllowPrintDocument; $parameters["allowFillForms"] = $AllowFillForms; $parameters["allowModifyDocument"] = $AllowModifyDocument; $parameters["allowContentExtraction"] = $AllowContentExtraction; $parameters["allowModifyAnnotations"] = $AllowModifyAnnotations; $parameters["printQuality"] = $PrintQuality; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the result file echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# Auto-rotate Pages with AI Source: https://developer.pdf.co/api/pdf-rotate/auto Automatically detect and fix page rotation in scanned PDFs using AI-based text analysis, with a configurable OCR language parameter. **Try it live:** [Auto-rotate Pages with AI → API Tester](/api-tester/pdf-rotate/auto) — send a real request from your browser. ## `POST /v1/pdf/edit/rotate/auto` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `fileSize` | integer | Size of the optimized PDF file in bytes | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf", "name": "result.pdf" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/HQ86WA7MFED1Q7C843NDSKVRW5AFVTMP/result.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEFwaDD6ndEhlfId4KouQ8yKCASr8amowIV2tLAi%2BhjnlVi%2FNYjf8ZJ3MqgKWsYVm5dQ8fQx7hceGdmtqhB6OH8t9xdacbMEcoMIpQr1BcSSfu2ZFfGFBDaHNpaSTfPXhkNnQaZFOi5KFozZiPBP9xPoSCV%2Fj%2BLIrDsOF%2Fb89i1Nd4OJFoXnfhjf03ZHJ%2BNCQEC%2BbePsovsjNlQYyKIV1qb1wGdwgJ%2ByibJ5x%2BQGoG4x2ebnEGTQkKBf4zobYT9Uv6FVQ%2FJg%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHB3D5EZ56/20220623/us-west-2/s3/aws4_request&X-Amz-Date=20220623T141005Z&X-Amz-SignedHeaders=host&X-Amz-Signature=351ad45980cad0d3f634f7b784d5445b9c4e1ad2244912f56ea065209d73bb30", "fileSize": 455115, "pageCount": 3, "error": false, "status": 200, "name": "result.pdf", "credits": 84, "duration": 6002, "remainingCredits": 98194629 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/rotate/auto' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf", "name": "result.pdf" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Rotate PDF` API endpoint var queryPath = `/v1/pdf/edit/rotate/auto`; // JSON payload for api request var jsonPayload = JSON.stringify({ url: SourceFileUrl, name: path.basename(DestinationFile), angle: Angle, pages: Pages }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): rotatePDF(SourceFileURL, DestinationFile) def rotatePDF(uploadedFileUrl, destinationFile): """Auto Rotate PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["url"] = uploadedFileUrl parameters["name"] = os.path.basename(destinationFile) # Prepare URL for 'Auto Rotate PDF' API request url = "{}/pdf/edit/rotate/auto".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Auto Rotate PDF` API call string url = "https://api.pdf.co/v1/pdf/edit/rotate/auto"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileUrl); parameters.Add("name", Path.GetFileName(DestinationFile)); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Rotate PDF` API call String query = "https://api.pdf.co/v1/pdf/edit/rotate/auto"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"url\": \"%s\", \"name\": \"%s\"}", SourceFileUrl, DestinationFile.getFileName()); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Optimize PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Rotate Selected Pages Source: https://developer.pdf.co/api/pdf-rotate/basic Rotates selected pages inside a PDF file. **Try it live:** [Rotate Selected Pages → API Tester](/api-tester/pdf-rotate/basic) — send a real request from your browser. ## `POST /v1/pdf/edit/rotate` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `angle` | integer | *No* | `0` | Contains the angle of rotation for the PDF pages. The available angles are: `0`, `90`, `180`, `270`. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `fileSize` | integer | Size of the optimized PDF file in bytes | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf", "name": "result.pdf" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/2EK4QYIZU1XUEUH853VTSK47NPLXCUYX/result.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEFsaDC5Vfgoi83YzdW9HXiKCAYBVHK096wqoUyu8Ckq8jEhV1DBv9VzHY1EcPWvfG3L2YrFa8QC5ZMr3UhFEn4%2B%2B2u6e%2FcdZd%2FXbdVaI45yNE%2Btz28UHMVxCQUClj9kCHrMyJ4W1%2BlnDgLi9JUHt7SkIvV9Lj7GLDBOXy22KCND86HdtPg0uT%2FNQtcjJm%2F34cISImKYov63NlQYyKG%2BEO%2FQLP%2BzJuugBdSLcKUOTL52dnc1l82ye1u5kYvTlbfPMdisU1tY%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHESTCUXXF/20220623/us-west-2/s3/aws4_request&X-Amz-Date=20220623T135902Z&X-Amz-SignedHeaders=host&X-Amz-Signature=78f2fe7f22ebd1fc2c329a855c7582f9deb09c0c5045281b9b2420a5afb792cf", "fileSize": 1064923, "pageCount": 4, "error": false, "status": 200, "name": "result.pdf", "credits": 28, "duration": 245, "remainingCredits": 98194839 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/rotate' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf", "name": "result.pdf", "angle": 90, "pages": "0-2,4" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf"; // Angle in degrees. Supported values are 90, 180, 270. const Angle = 90; // Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty const Pages = "0-2,4"; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Rotate PDF` API endpoint var queryPath = `/v1/pdf/edit/rotate`; // JSON payload for api request var jsonPayload = JSON.stringify({ url: SourceFileUrl, name: path.basename(DestinationFile), angle: Angle, pages: Pages, async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); function checkIfJobIsCompleted(jobId, resultFileUrl) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl); }, 3000); } else if (data.status == "success") { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf" # Angle in degrees. Supported values are 90, 180, 270. Angle = 90 # Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty Pages = "0-2,4" # Destination PDF file name DestinationFile = ".\\result.pdf" # (!) Make asynchronous job Async = True def main(args = None): rotatePDF(SourceFileURL, DestinationFile) def rotatePDF(uploadedFileUrl, destinationFile): """Rotate PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["url"] = uploadedFileUrl parameters["name"] = os.path.basename(destinationFile) parameters["angle"] = Angle parameters["pages"] = Pages parameters["async"] = Async # Prepare URL for 'Rotate PDF' API request url = "{}/pdf/edit/rotate".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # URL of the result file resultFileUrl = json["url"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.IO; using System.Net; using Newtonsoft.Json.Linq; using System.Threading; using System.Collections.Generic; using Newtonsoft.Json; // Cloud API asynchronous "Rotate PDF" job example. namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf"; // Angle in degrees. Supported values are 90, 180, 270. const int Angle = 90; // Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty const string Pages = "0-2,4"; // Destination PDF file name const string DestinationFile = @".\result.pdf"; // (!) Make asynchronous job const bool Async = true; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Rotate PDF` API call string url = "https://api.pdf.co/v1/pdf/edit/rotate"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("url", SourceFileUrl); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("angle", Angle); parameters.Add("pages", Pages); parameters.Add("async", Async); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Asynchronous job ID string jobId = json["jobId"].ToString(); // URL of generated PDF file that will available after the job completion string resultFileUrl = json["url"].ToString(); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status); if (status == "success") { // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); break; } else if (status == "working") { // Pause for a few seconds Thread.Sleep(3000); } else { Console.WriteLine(status); break; } } while (true); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } static string CheckJobStatus(string jobId) { using (WebClient webClient = new WebClient()) { // Set API Key webClient.Headers.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; string response = webClient.DownloadString(url); JObject json = JObject.Parse(response); return Convert.ToString(json["status"]); } } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf"; // Angle in degrees. Supported values are 90, 180, 270. final static Integer Angle = 90; // Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty final static String Pages = "0-2,4"; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Rotate PDF` API call String query = "https://api.pdf.co/v1/pdf/edit/rotate"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"url\": \"%s\", \"name\": \"%s\", \"angle\": \"%d\", \"pages\": \"%s\"}", SourceFileUrl, DestinationFile.getFileName(), Angle, Pages); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Optimize PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF Search and Delete Text Source: https://developer.pdf.co/api/pdf-search-text-and-delete Delete text from the PDF document with search strings. **Try it live:** [PDF Search and Delete Text → API Tester](/api-tester/pdf-search-text-and-delete) — send a real request from your browser. ## `POST /v1/pdf/edit/delete-text` When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `searchString` | string | *Yes* | - | Single string to search and delete. Provide either `searchString` or `searchStrings`. Text to search can support regular expressions if you set the `regex` param to true. | | `searchStrings` | array\[string] | *Yes* | - | Array of strings to search and delete. Provide either `searchString` or `searchStrings`. | | `redactions` | array | *No* | - | A list of redaction objects, each defining a rectangular area to remove or cover. Each redaction object includes the following fields: | |     `page` | integer | *Yes* | - | The zero-based index of the page where the redaction should be applied. For example, 0 = first page. | |     `x` | float | *Yes* | - | The X-coordinate of the top-left corner of the redaction box, in PDF units (points). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |     `y` | float | *Yes* | - | The Y-coordinate of the top-left corner of the redaction box, in PDF units (points). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |     `width` | float | *Yes* | - | The width of the redaction area in points. [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | |     `height` | float | *Yes* | - | The height of the redaction area in points. [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. | | `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. | | `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. | | `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). | | `MakeUnsearchable` | boolean | *No* | `false` | If `true`, the output PDF is made unsearchable by replacing all pages with images. If `false`, the document remains searchable with only the specified text removed. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. If not specified, the default configuration processes all pages. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `removeTextUnderPatch` | boolean | *No* | `true` | Controls whether to remove text under the patch or not | |     `usepatch` | boolean | *No* | `false` | Controls whether to use a patch or not | |     `patchColor` | string | *No* | `#000000` | Controls the color of the patch | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Showing Redacted Text By default when we delete text using [post-tag-pdf-edit-delete-text](/api/pdf-search-text-and-delete) it will simply remove text leaving a space where the text was. In the case where you need to blackout deleted text it can be acheived using following `profiles` parameters. * Set `UsePatch` parameter to `true`. * Set `PatchColor` parameter to color we want to use for redacting in `hex` format. For example: `'PatchColor': '#000000'`. In case we want to only blackout text, but *not remove it* so that we can still copy it, we can do so using `RemoveTextUnderPatch` parameter and set it to `false`. If `RemoveTextUnderPatch` is set to `false` then a user could still copy the text making the redaction less secure than you might require! ``` { "profiles": "{'UsePatch': true, 'PatchColor': '#000000', 'RemoveTextUnderPatch': true}" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "name": "pdfWithTextDeleted", "caseSensitive": "false", "searchString": "Invoice", "replacementLimit": 0, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/ZOSEQZFNVCYLD5N5CJFVIYQKBVLR8OKD/pdfWithTextDeleted.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzECYaDKOO4WmO5C5shyOYYSKCAVsAo6VkB5HQjTBd9dMlJujQdEkPfNdPeLfq2mF54s2ESZBmIAJ5UgDUo3J9R475CCS4M3nuuo%2FSJwRy5gNiJdb1ZY0uCtP87x83nH%2B%2BSDu5JK%2F%2BEOrd3MREt8KE3BsQOrv%2FKMdnK%2BT5nJ2x2hC87vHue%2FudY7%2FWX54vx4tfFobEyhEozLbPnwYyKOdEsYYWH7e8tm7XV4UeKxCoKMaXSEPvOod80hR62qXnEI42fOsON3M%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHLUVIAIPX/20230220/us-west-2/s3/aws4_request&X-Amz-Date=20230220T205521Z&X-Amz-SignedHeaders=host&X-Amz-Signature=9f79c1a30d4f373e495e735e908375dad2ae6dcafcee761a477748c2b8298605", "pageCount": 1, "error": false, "status": 200, "name": "pdfWithTextDeleted.pdf", "credits": 21, "duration": 189, "remainingCredits": 96235635 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/delete-text' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "name": "pdfWithTextDeleted", "caseSensitive": "false", "searchString": "Invoice", "replacementLimit": 0, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Delete Text from PDF` API endpoint var queryPath = `/v1/pdf/edit/delete-text`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, url: SourceFileUrl, searchString: 'conspicuous' }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): deleteTextFromPdf(SourceFileURL, DestinationFile) def deleteTextFromPdf(uploadedFileUrl, destinationFile): """Delete Text from PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["url"] = uploadedFileUrl parameters["searchString"] = "conspicuous" # Prepare URL for 'Delete Text from PDF' API request url = "{}/pdf/edit/delete-text".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination PDF file name const string DestinationFile = @".\result.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("url", SourceFileUrl); parameters.Add("searchString", "conspicuous"); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // URL of `Delete Text from PDF` API call string url = "https://api.pdf.co/v1/pdf/edit/delete-text"; try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated PDF file string resultFileUrl = json["url"].ToString(); // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "**********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Delete Text from PDF` API call String query = "https://api.pdf.co/v1/pdf/edit/delete-text"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"conspicuous\"}", DestinationFile.getFileName(), Password, SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Delete Text from PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Search and Replace with Image Source: https://developer.pdf.co/api/pdf-search-text-and-replace/image Modify a PDF file by searching for specific text and replacing it with an image. **Try it live:** [Search and Replace with Image → API Tester](/api-tester/pdf-search-text-and-replace/image) — send a real request from your browser. ## `POST /v1/pdf/edit/replace-text-with-image` When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. | | `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. | | `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `searchString` | string | *Yes* | - | Single text replacement. Word or phrase to be replaced. Text to search can support regular expressions if you set the `regex` param to true. | | `replaceImage` | string | *Yes* | - | Image URL or datauri to be inserted in the document | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `AutoCropImages` | boolean | *No* | - | Controls whether to crop empty space around an inserted image. See [Crop Empty Space Around Images](#crop-empty-space-around-images) for more information. | ### Crop Empty Space Around Images If you require to crop empty space around an inserted image use the following: ```json theme={null} { "profiles": "{'AutoCropImages': true}" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "searchString": "Your Company Name", "caseSensitive": false, "replaceImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png", "pages": "0", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/7ea2b532988742508906cff59be0180e/sample.pdf", "pageCount": 1, "error": false, "status": 200, "name": "sample.pdf", "remainingCredits": 99150679, "credits": 77 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/replace-text-with-image' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "searchString": "Your Company Name", "caseSensitive": false, "replaceImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png", "pages": "0", "async": false }' ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-2.pdf" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination PDF file name DestinationFile = ".\\result.pdf" # (!) Make asynchronous job Async = True def main(args = None): replaceImageFromPdf(SourceFileURL, DestinationFile) def replaceImageFromPdf(uploadedFileUrl, destinationFile): """Replace Text With Image from PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co parameters = {} parameters["async"] = Async parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["url"] = uploadedFileUrl parameters["searchString"] = "[CLIENT-SIGNATURE]" parameters["caseSensitive"] = True parameters["replaceImage"] = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/john-doe-signature-image.png" # Prepare URL for 'Replace Text With Image from PDF' API request url = "{}/pdf/edit/replace-text-with-image".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # URL of the result file resultFileUrl = json["url"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.IO; using System.Net; using Newtonsoft.Json.Linq; using System.Threading; using System.Collections.Generic; using Newtonsoft.Json; // Cloud API asynchronous "Replace Text With Image from PDF" job example. namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "*****************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-2.pdf"; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination PDF file name const string DestinationFile = @".\result.pdf"; // (!) Make asynchronous job const bool Async = true; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `Replace Text With Image from PDF` API call string url = "https://api.pdf.co/v1/pdf/edit/replace-text-with-image"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("url", SourceFileUrl); parameters.Add("async", Async); parameters.Add("searchString", "[CLIENT-SIGNATURE]"); parameters.Add("caseSensitive", true); parameters.Add("replaceImage", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/john-doe-signature-image.png"); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Asynchronous job ID string jobId = json["jobId"].ToString(); // URL of generated PDF file that will available after the job completion string resultFileUrl = json["url"].ToString(); // Check the job status in a loop. // If you don't want to pause the main thread you can rework the code // to use a separate thread for the status checking and completion. do { string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status); if (status == "success") { // Download PDF file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile); break; } else if (status == "working") { // Pause for a few seconds Thread.Sleep(3000); } else { Console.WriteLine(status); break; } } while (true); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } static string CheckJobStatus(string jobId) { using (WebClient webClient = new WebClient()) { // Set API Key webClient.Headers.Add("x-api-key", API_KEY); string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId; string response = webClient.DownloadString(url); JObject json = JObject.Parse(response); return Convert.ToString(json["status"]); } } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Replace Text With Image from PDF` API call String query = "https://api.pdf.co/v1/pdf/edit/replace-text-with-image"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"/creativecommons.org/licenses/by-sa/3.0/\", \"replaceImage\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png\"}", DestinationFile.getFileName(), Password, SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Replace Text With Image from PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Search and Replace with Text Source: https://developer.pdf.co/api/pdf-search-text-and-replace/text Replaces text in a PDF file with a new text. **Try it live:** [Search and Replace with Text → API Tester](/api-tester/pdf-search-text-and-replace/text) — send a real request from your browser. ## `POST /v1/pdf/edit/replace-text` When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `searchString` | string | *Yes* | - | Single string to search. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. Text to search can support regular expressions if you set the `regex` param to true. | | `replaceString` | string | *Yes* | - | Single replacement string. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. | | `searchStrings` | array\[string] | *Yes* | - | Array of strings to search. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. | | `replaceStrings` | array\[string] | *Yes* | - | Array of replacement strings. Must match the order of `searchStrings`. | | `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. | | `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. | | `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `AutoCropImages` | boolean | *No* | `false` | If you require to crop empty space around an inserted image use the following: `profiles": { 'AutoCropImages': true }` | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `YAdjustmentForReplacementText` | integer | *No* | - | Adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. See [Adjust Text Alignment](#adjust-text-alignment) for more details. | ### Adjust Text Alignment Users may have encountered an issue when using this API endpoint to replace text in a **PDF** document. The replaced text might appear slightly higher than the original text or the surrounding text, causing alignment issues. To fix this issue, we have added a new parameter called `YAdjustmentForReplacementText` in the `profiles` parameter of the API request. This parameter allows you to adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. Negative values for this parameter move text up, positive values move text down. Here’s an example of how to use the `YAdjustmentForReplacementText` parameter. In this example API request, the `YAdjustmentForReplacementText` parameter has been set to `-1`, which moves the replaced text `1` unit up vertically, resulting in better alignment with the original text. ```json theme={null} { "profiles": "{'YAdjustmentForReplacementText': '-1'}" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-1.pdf", "searchStrings": [ "[CLIENT-NAME]", "[CLIENT-COMPANY]" ], "replaceStrings": [ "John Doe", "Skynet 3000" ], "caseSensitive": true, "replacementLimit": 1, "pages": "", "password": "", "name": "finalFile", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/e79f0b9c82984740973ca670d7c93cad/finalFile.pdf", "pageCount": 1, "error": false, "status": 200, "name": "finalFile.pdf", "remainingCredits": 99089875, "credits": 21 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/replace-text' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "name": "pdfWithTextReplaced", "caseSensitive": "false", "searchString": "Your Company Name", "replaceString": "Acme ltd.", "replacementLimit": 0, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination PDF file name const DestinationFile = "./result.pdf"; // Prepare request to `Replace Text from PDF` API endpoint var queryPath = `/v1/pdf/edit/replace-text`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), password: Password, url: SourceFileUrl, searchString: 'Your Company Name', replaceString: 'XYZ LLC' }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download PDF file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Direct URL of source PDF file. # You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination PDF file name DestinationFile = ".\\result.pdf" def main(args = None): replaceStringFromPdf(SourceFileURL, DestinationFile) def replaceStringFromPdf(uploadedFileUrl, destinationFile): """Replace Text from PDF using PDF.co Web API""" # Prepare requests params as JSON parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["url"] = uploadedFileUrl parameters["searchString"] = "Your Company Name" parameters["replaceString"] = "XYZ LLC" # Prepare URL for 'Replace Text from PDF' API request url = "{}/pdf/edit/replace-text".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") if __name__ == '__main__': main() ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination PDF file name final static Path DestinationFile = Paths.get(".\\result.pdf"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `Replace Text from PDF` API call String query = "https://api.pdf.co/v1/pdf/edit/replace-text"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"Your Company Name\", \"replaceString\": \"XYZ LLC\"}", DestinationFile.getFileName(), Password, SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated PDF file String resultFileUrl = json.get("url").getAsString(); // Download PDF file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} Cloud API asynchronous "Replace Text from PDF" job example (allows to avoid timeout errors). " . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# Split PDF Source: https://developer.pdf.co/api/pdf-split/by-pages Split a PDF into multiple files by specifying page numbers or page ranges to keep in each output. **Try it live:** [Split PDF → API Tester](/api-tester/pdf-split/by-pages) — send a real request from your browser. ## `POST /v1/pdf/split` When splitting a document the pages parameter controls which `pages` to split out into individual documents. The page limit should not exceed the number of pages in the document - for example, you cannot split a 100 page document into 200 individual documents, however you can split it into 100 individual documents. The `pages` parameter is 1-based, meaning the first page is `1` and not `0`. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify pages as comma-separated page numbers and ranges to process (e.g. "1, 2, 5-10" or "3-" for page 3 to the end). The first-page index is 1. Use "!" before a number for inverted page numbers (e.g. "!1" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `fixed_output_filename` | boolean | *No* | `false` | Defines whether the page range is appended to the output filename. When set to `true`, the output filename remains exactly as specified in the `name` parameter. When set to `false` (default), the filename automatically includes the page range of the extracted pages for each output file. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf", "pages": "1-2,3-", "inline": true, "name": "result.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/1e9a7f2c46834160903276716424382b/result_page1-2.pdf", "https://pdf-temp-files.s3.amazonaws.com/c976b9f89a2e460786a3d5c0deeeef67/result_page3-4.pdf" ], "pageCount": 4, "error": false, "status": 200, "name": "result.pdf", "remainingCredits": 98441 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/split' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf", "pages": "1-2,3-", "inline": true, "name": "result.pdf", "async": false }' ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'. Pages = "1-2,3-" # (!) Make asynchronous job Async = True def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): splitPDF(uploadedFileUrl) def splitPDF(uploadedFileUrl): """Split PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["async"] = Async parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'Split PDF' API request url = "{}/pdf/split".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # URL of the result file resultFilePlaceholder = json["url"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": resJsonImgFiles = requests.get(resultFilePlaceholder) # Download generated PNG files part = 1 for resultFileUrl in resJsonImgFiles.json(): # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.pdf" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file to split const string SourceFile = @".\sample.pdf"; // Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'. const string Pages = "1-2,3-"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. SPLIT UPLOADED PDF // URL for `Split PDF` API call var url = "https://api.pdf.co/v1/pdf/split"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated PDF files int part = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\part{0}.pdf", part); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); part++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file to split final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'. final static String Pages = "1-2,3-"; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. SPLIT UPLOADED PDF SplitPdf(webClient, API_KEY, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void SplitPdf(OkHttpClient webClient, String apiKey, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `Split PDF` API call String query = "https://api.pdf.co/v1/pdf/split"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"pages\": \"%s\", \"url\": \"%s\"}", pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated PDF files JsonArray urls = json.get("urls").getAsJsonArray(); int part = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\part%s.pdf", part); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Splitted part saved as \"%s\".", localFileName)); part++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF Splitting Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function SplitPdf($apiKey, $fileUrl, $pages) { // Prepare URL for `Split PDF` API call $url = "https://api.pdf.co/v1/pdf/split"; // Prepare requests params $parameters = array(); $parameters["name"] = "part.pdf"; $parameters["url"] = $fileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // Display links to splitted parts $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display request error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# Split PDF by Text Search Source: https://developer.pdf.co/api/pdf-split/by-text-search-or-barcode Split a PDF into multiple files at every page that matches a text search or barcode pattern. **Try it live:** [Split PDF by Text Search → API Tester](/api-tester/pdf-split/by-text-search-or-barcode) — send a real request from your browser. ## `POST /v1/pdf/split2` This endpoint decides where to split using `searchString` only. Split points are every page that matches the search, so there is no page-selection parameter to configure. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `searchString` | string | *Yes* | - | Text to search for on pages. Must be a string. | | `regexSearch` | boolean | *No* | `false` | Set to true to enable regular expression search for the `searchString(s)` parameter. | | `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `excludeKeyPages` | boolean | *No* | `false` | Set to true to exclude pages where the searchString text was found. | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ### searchString Text to search for on pages. Must be a string. To search for a barcode use the following macros string: `[[barcode: ]]`. To search for barcode type without analyzing its value, use this notation instead: `[[barcode:]].` Example #1, split by QR code: "searchString": "\[\[barcode:qrcode]]". Example #2, split by QR code with value: "searchString": "\[\[barcode:qrcode pdfco]]". Example #3, split by QR code with value search with regex: "searchString": "\[\[barcode:qrcode /pdf.co/]]". Example #4, split by QR code or datamatrix with value search with regex: "searchString": "\[\[barcode:qrcode,datamatrix /pdf.co/]]". ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/split_by_barcode.pdf", "searchString": "[[barcode:qrcode,datamatrix /pdf\\.co/]]", "excludeKeyPages": true, "regexSearch": false, "caseSensitive": false, "inline": true, "name": "output-split-by-barcode", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.us-west-2.amazonaws.com/A2WX2GR0PX4818EIKW96VR3BZTK5FWT2/output-split-by-barcode_page1.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=8241ad05ecb5555cbbd4998b5c334104f2849bf4177384e86fbb5cc5d7e81ce8", "https://pdf-temp-files.s3.us-west-2.amazonaws.com/B6Z9J274GZ5BK5QYK547ST4T5WF61LNQ/output-split-by-barcode_page3-5.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=94764cfb37819f2a4885ba064dd1ae20f38f42d6bc6c1a208010637fca74a591", "https://pdf-temp-files.s3.us-west-2.amazonaws.com/XT5TD1BDBFDNKX0LM6N5GLFLOAF1UC0Y/output-split-by-barcode_page7-9.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=0a7c90a05fd159659451d29273284fbf422d34bd204c07fbc9abdf7a36a84294" ], "pageCount": 10, "error": false, "status": 200, "name": "output-split-by-barcode.pdf", "credits": 350, "duration": 4456, "remainingCredits": 98221710 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/split2' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/split_by_barcode.pdf", "searchString": "[[barcode:qrcode,datamatrix /pdf\\.co/]]", "excludeKeyPages": true, "regexSearch": false, "caseSensitive": false, "inline": true, "name": "output-split-by-barcode", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file to split const SourceFile = "./sample.pdf"; // Split Search String const SplitText = "invoice number"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. SPLIT UPLOADED PDF splitPdf(API_KEY, uploadedFileUrl, SplitText); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + err); } }); }); }); } function splitPdf(apiKey, uploadedFileUrl, splitText) { // Prepare request to `Split PDF By Text` API endpoint var queryPath = `/v1/pdf/split2`; // JSON payload for api request var jsonPayload = JSON.stringify({ searchString: splitText, url: uploadedFileUrl, async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { console.log(`Job #${data.jobId} has been created!`); checkIfJobIsCompleted(data.jobId, data.url); } else { // Service reported error console.log("splitPdf(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("splitPdf(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } function checkIfJobIsCompleted(jobId, resultFileUrlJson) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrlJson) }, 3000); } else if (data.status == "success") { request({ method: 'GET', uri: resultFileUrlJson, gzip: true }, function (error, response, body) { // Parse JSON response let respJsonFileArray = JSON.parse(body); let part = 1; respJsonFileArray.forEach((url) => { var localFileName = `./part${part}.pdf`; var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated PDF file saved as "${localFileName} file."`); }); }); part++; }, this); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests import time import datetime # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Split by Text SplitText = "invoice number" # (!) Make asynchronous job Async = True def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): splitPDF(uploadedFileUrl) def splitPDF(uploadedFileUrl): """Split PDF using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["async"] = Async parameters["searchString"] = SplitText parameters["url"] = uploadedFileUrl # Prepare URL for 'Split PDF By Text' API request url = "{}/pdf/split2".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Asynchronous job ID jobId = json["jobId"] # URL of the result file resultFilePlaceholder = json["url"] # Check the job status in a loop. # If you don't want to pause the main thread you can rework the code # to use a separate thread for the status checking and completion. while True: status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success". # Display timestamp and status (for demo purposes) print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status) if status == "success": resJsonImgFiles = requests.get(resultFilePlaceholder) # Download generated files part = 1 for resultFileUrl in resJsonImgFiles.json(): # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.pdf" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 break elif status == "working": # Pause for a few seconds time.sleep(3) else: print(status) break else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def checkJobStatus(jobId): """Checks server job status""" url = f"{BASE_URL}/job/check?jobid={jobId}" response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() return json["status"] else: print(f"Request error: {response.status_code} {response.reason}") return None def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file to split const string SourceFile = @".\sample.pdf"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. SPLIT UPLOADED PDF By Text // URL for `Split PDF By Text` API call var url = "https://api.pdf.co/v1/pdf/split2"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("searchString", "invoice number"); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated PDF files int part = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\part{0}.pdf", part); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); part++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file to split final static Path SourceFile = Paths.get(".\\sample.pdf"); // Split By Text final static String SplitText = "invoice number"; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. SPLIT UPLOADED PDF SplitPdf(webClient, API_KEY, SplitText, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void SplitPdf(OkHttpClient webClient, String apiKey, String splitText, String uploadedFileUrl) throws IOException { // Prepare URL for `Split PDF By Text` API call String query = "https://api.pdf.co/v1/pdf/split2"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"searchString\": \"%s\", \"url\": \"%s\"}", splitText, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated PDF files JsonArray urls = json.get("urls").getAsJsonArray(); int part = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\part%s.pdf", part); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Splitted part saved as \"%s\".", localFileName)); part++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF Splitting Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function SplitPdf($apiKey, $fileUrl, $splitText) { // Prepare URL for `Split PDF By Text` API call $url = "https://api.pdf.co/v1/pdf/split2"; // Prepare requests params $parameters = array(); $parameters["name"] = "part.pdf"; $parameters["url"] = $fileUrl; $parameters["searchString"] = $splitText; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // Display links to splitted parts $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display request error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to CSV Source: https://developer.pdf.co/api/pdf-to-csv Convert PDF and scanned images into CSV representation with layout, columns, rows, and tables. **Try it live:** [PDF to CSV → API Tester](/api-tester/pdf-to-csv) — send a real request from your browser. ## `POST /v1/pdf/convert/to/csv` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. | |     `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | `None` | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file. ### Column Detection Mode This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data. Set the options for your column detection via the following `profiles` parameters: `ColumnDetectionMode` - available values: * `ContentGroupsAndBorders` (default, no need to specify) * `ContentGroups` * `Borders` * `BorderedTables` * `ContentGroupsAI` ```json theme={null} { "profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }" } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "ExtractShadowLikeText": false, "OCRMode": "Auto", "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `body` | string | Stringified CSV content | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf", "lang": "eng", "inline": "true", "unwrap": "", "pages": "0-", "rect": "", "async": "false", "name": "result.csv", "password": "", "lineGrouping": "", "profiles": "" } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": "\"Your Company Name\",\"\",\"\",\"\",\r\n\"Your Address\",\"\",\"\",\"\",\r\n\"City, State Zip\",\"\",\"\",\"\",\r\n\"\",\"\",\"\",\"Invoice No. 123456\",\r\n\"\",\"\",\"\",\"Invoice Date 01/01/2016\",\r\n\"Client Name\",\"\",\"\",\"\",\r\n\"Address\",\"\",\"\",\"\",\r\n\"City, State Zip\",\"\",\"\",\"\",\r\n\"Notes\",\"\",\"\",\"\",\r\n\"Item\",\"Quantity\",\"Price\",\"Total\",\r\n\"Item 1\",\"1\",\"40.00\",\"40.00\",\r\n\"Item 2\",\"2\",\"30.00\",\"60.00\",\r\n\"Item 3\",\"3\",\"20.00\",\"60.00\",\r\n\"Item 4\",\"4\",\"10.00\",\"40.00\",\r\n\"\",\"\",\"TOTAL\",\"200.00\",\r\n", "pageCount": 2, "error": false, "status": 200, "name": "result.csv", "remainingCredits": 616411, "credits": 56 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/csv' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf", "lang": "eng", "inline": "true", "unwrap": "", "pages": "0-", "rect": "", "async": "false", "name": "result.csv", "password": "", "lineGrouping": "", "profiles": "" }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "*********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination CSV file name const DestinationFile = "./result.csv"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO CSV convertPdfToCsv(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToCsv(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To CSV` API endpoint var queryPath = `/v1/pdf/convert/to/csv`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl, async: true }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Job #${data.jobId} has been created!`); if (data.error == false) { checkIfJobIsCompleted(data.jobId, data.url, destinationFile); } else { // Service reported error console.log("convertPdfToCsv(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToCsv(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) { let queryPath = `/v1/job/check`; // JSON payload for api request let jsonPayload = JSON.stringify({ jobid: jobId }); let reqOptions = { host: "api.pdf.co", path: queryPath, method: "POST", headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`); if (data.status == "working") { // Check again after 3 seconds setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000); } else if (data.status == "success") { // Download CSV file var file = fs.createWriteStream(destinationFile); https.get(resultFileUrl, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated CSV file saved as "${destinationFile}" file.`); }); }); } else { console.log(`Operation ended with status: "${data.status}".`); } }) }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination CSV file name DestinationFile = ".\\result.csv" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToCSV(uploadedFileUrl, DestinationFile) def convertPdfToCSV(uploadedFileUrl, destinationFile): """Converts PDF To CSV using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To CSV' API request url = "{}/pdf/convert/to/csv".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "**************************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination CSV file name const string DestinationFile = @".\result.csv"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["status"].ToString() != "error") { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO CSV // URL for `PDF To CSV` API call var url = "https://api.pdf.co/v1/pdf/convert/to/csv"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["status"].ToString() != "error") { // Get URL of generated CSV file string resultFileUrl = json["url"].ToString(); // Download CSV file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated CSV file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination CSV file name final static Path DestinationFile = Paths.get(".\\result.csv"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); String status = json.get("status").getAsString(); if (!status.equals("error")) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO CSV PdfToCsv(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToCsv(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To CSV` API call String query = "https://api.pdf.co/v1/pdf/convert/to/csv"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); String status = json.get("status").getAsString(); if (!status.equals("error")) { // Get URL of generated CSV file String resultFileUrl = json.get("url").getAsString(); // Download CSV file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated CSV file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To CSV Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractCSV($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/csv"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to XLS Source: https://developer.pdf.co/api/pdf-to-excel/xls Convert PDF to Excel(.xls) with layout and fonts preserved. **Try it live:** [PDF to XLS → API Tester](/api-tester/pdf-to-excel/xls) — send a real request from your browser. ## `POST /v1/pdf/convert/to/xls` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file. ### Column Detection Mode This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data. Set the options for your column detection via the following `profiles` parameters: `ColumnDetectionMode` - available values: * `ContentGroupsAndBorders` (default, no need to specify) * `ContentGroups` * `Borders` * `BorderedTables` * `ContentGroupsAI` ```json theme={null} { "profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }" } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "ExtractShadowLikeText": false, "OCRMode": "Auto", "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/60c6b9f50280495a9567f73a0a394252/sample.xlsx", "pageCount": 1, "error": false, "status": 200, "name": "sample.xlsx", "remainingCredits": 60568 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination XLS file name const DestinationFile = "./result.xls"; // Prepare request to `PDF To XLS` API endpoint var queryPath = `/v1/pdf/convert/to/xls`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download XLS file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated XLS file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination XLS file name const string DestinationFile = @".\result.xls"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // URL for `PDF To XLS` API call string url = "https://api.pdf.co/v1/pdf/convert/to/xls"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", SourceFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); try { // Execute POST request with JSON payload string response = webClient.UploadString(url, jsonPayload); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated XLS file string resultFileUrl = json["url"].ToString(); // Download XLS file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated XLS file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination XLS file name final static Path DestinationFile = Paths.get(".\\result.xls"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // Prepare URL for `PDF To XLS` API call String query = "https://api.pdf.co/v1/pdf/convert/to/xls"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", DestinationFile.getFileName(), Password, Pages, SourceFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated XLS file String resultFileUrl = json.get("url").getAsString(); // Download XLS file downloadFile(webClient, resultFileUrl, DestinationFile.toFile()); System.out.printf("Generated XLS file saved as \"%s\" file.", DestinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } PDF To Excel Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractExcel($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/xlsx"; // (!) If you need the old XLS format use `https://api.pdf.co/v1/pdf/convert/to/xls` endpoint, // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; $parameters["async"] = true; // (!) Make asynchronous job // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // URL of generated XLSX file that will available after the job completion $resultFileUrl = $json["url"]; // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF to XLSX Source: https://developer.pdf.co/api/pdf-to-excel/xlsx Convert PDF to Excel(.xlsx) with layout and fonts preserved. **Try it live:** [PDF to XLSX → API Tester](/api-tester/pdf-to-excel/xlsx) — send a real request from your browser. ## `POST /v1/pdf/convert/to/xlsx` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file. ### Column Detection Mode This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data. Set the options for your column detection via the following `profiles` parameters: `ColumnDetectionMode` - available values: * `ContentGroupsAndBorders` (default, no need to specify) * `ContentGroups` * `Borders` * `BorderedTables` * `ContentGroupsAI` ```json theme={null} { "profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }" } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "ExtractShadowLikeText": false, "OCRMode": "Auto", "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/60c6b9f50280495a9567f73a0a394252/sample.xlsx", "pageCount": 1, "error": false, "status": 200, "name": "sample.xlsx", "remainingCredits": 60568 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/xlsx?=' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Direct URL of source PDF file. // You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/ const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination XLSX file name const DestinationFile = "./result.xlsx"; // Prepare request to `PDF To XLSX` API endpoint var queryPath = `/v1/pdf/convert/to/xlsx`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { // Parse JSON response var data = JSON.parse(d); if (data.error == false) { // Download XLSX file var file = fs.createWriteStream(DestinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated XLSX file saved as "${DestinationFile}" file.`); }); }); } else { // Service reported error console.log(data.message); } }); }).on("error", (e) => { // Request error console.log(e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination Excel file name DestinationFile = ".\\result.xlsx" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToExcel(uploadedFileUrl, DestinationFile) def convertPdfToExcel(uploadedFileUrl, destinationFile): """Converts PDF To Excel using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/ parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Xlsx' API request url = "{}/pdf/convert/to/xlsx".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination XLSX file name const string DestinationFile = @".\result.xlsx"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO XLSX // URL for `PDF To XLSX` API call var url = "https://api.pdf.co/v1/pdf/convert/to/xlsx"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated XLSX file string resultFileUrl = json["url"].ToString(); // Download XLSX file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated XLSX file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination XLSX file name final static Path DestinationFile = Paths.get(".\\result.xlsx"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO XLSX PdfToXlsx(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToXlsx(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To XLSX` API call String query = "https://api.pdf.co/v1/pdf/convert/to/xlsx"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated XLSX file String resultFileUrl = json.get("url").getAsString(); // Download XLSX file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated XLSX file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Excel Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractExcel($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/xlsx"; // (!) If you need the old XLS format use `https://api.pdf.co/v1/pdf/convert/to/xls` endpoint, // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; $parameters["async"] = true; // (!) Make asynchronous job // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { // URL of generated XLSX file that will available after the job completion $resultFileUrl = $json["url"]; // Asynchronous job ID $jobId = $json["jobId"]; // Check the job status in a loop do { $status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success". // Display timestamp and status (for demo purposes) echo "

" . date(DATE_RFC2822) . ": " . $status . "

"; if ($status == "success") { // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; break; } else if ($status == "working") { // Pause for a few seconds sleep(3); } else { echo $status . "
"; break; } } while (true); } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } function CheckJobStatus($jobId, $apiKey) { $status = null; // Create URL $url = "https://api.pdf.co/v1/job/check"; // Prepare requests params $parameters = array(); $parameters["jobid"] = $jobId; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $status = $json["status"]; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); return $status; } ?> ```
# PDF to HTML Source: https://developer.pdf.co/api/pdf-to-html Convert PDF and scanned images into HTML representation with text, fonts, images, vectors, formatting preserved. **Try it live:** [PDF to HTML → API Tester](/api-tester/pdf-to-html) — send a real request from your browser. ## `POST /v1/pdf/convert/to/html` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `OptimizeImages` | boolean | *No* | `true` | Some PDF may have high quality images used in the document and you may need to keep the quality of these images in the output HTML. By default PDF to HTML is optimizing images and you can easily turn it off. See [Control Image Quality](#control-image-quality) for more information. | |     `OutputPageWidth` | integer | *No* | `1024` | Control page width (in pixels) for output HTML. Height is calculated and used according to the original pdf pages ratio. See [Control Output Page Width](#control-output-page-width) for more information. | |     `AdditionalCssStyles` | string | *No* | \`\` | To inject CSS for layout options in your HTML. Example: `#canvas { zoom: 50%; }`. Scale the div that contains all generated HTML pages by 50%. See [Inject CSS](#inject-css) for more information. | |     `saveImages` | integer | *No* | - | Controls whether to save images in the output HTML. See [Disable Images](#disable-images) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Disable Images To turn off images output set the following profile: ```json theme={null} { "profiles": "{ 'saveImages': 0 }" } ``` ### Control Image Quality Some **PDF** may have high quality images used in the document and you may need to keep the quality of these images in the output **HTML**. By default [PDF to HTML](/api/pdf-to-html) is optimizing images and you can easily turn it off with the following profile: ```json theme={null} { "profiles": "{ 'OptimizeImages': false }" } ``` ### Control Output Page Width Control page width output as follows: ```json theme={null} { "profiles": "{ 'OutputPageWidth': 2048 }" } ``` ### Inject CSS To inject CSS for layout options in your HTML use the following: ```json theme={null} { "profiles": "{ 'AdditionalCssStyles': '#canvas { zoom: 50%; }' }" } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "ExtractShadowLikeText": false, "OCRMode": "Auto", "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf", "inline": false, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://pdf-temp-files.s3.amazonaws.com/a7a86e9f29f84f5180624bdec1facfc2/index.html", "pageCount": 1, "error": false, "status": 200, "name": "sample.html", "remainingCredits": 60110 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/html' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf", "inline": false, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination HTML file name const DestinationFile = "./result.html"; // Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design. const PlainHtml = false; // Set to `true` if your document has the column layout like a newspaper. const ColumnLayout = false; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO HTML convertPdfToHtml(API_KEY, uploadedFileUrl, Password, Pages, PlainHtml, ColumnLayout, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToHtml(apiKey, uploadedFileUrl, password, pages, plainHtml, columnLayout, destinationFile) { // Prepare request to `PDF To HTML` API endpoint var queryPath = `/v1/pdf/convert/to/html`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, simple: plainHtml, columns: columnLayout, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download HTML file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated HTML file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToHtml(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToHtml(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "***************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination Html file name DestinationFile = ".\\result.html" # Set to $true to get simplified HTML without CSS. Default is the rich HTML keeping the document design. PlainHtml = False # Set to $true if your document has the column layout like a newspaper. ColumnLayout = False def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToHtml(uploadedFileUrl, DestinationFile) def convertPdfToHtml(uploadedFileUrl, destinationFile): """Converts PDF To Html using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-html parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["simple"] = PlainHtml parameters["columns"] = ColumnLayout parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Html' API request url = "{}/pdf/convert/to/html".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination HTML file name const string DestinationFile = @".\result.html"; // Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design. const bool PlainHtml = false; // Set to `true` if your document has the column layout like a newspaper. const bool ColumnLayout = false; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have the direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO HTML // URL for `PDF To HTML` API call var url = "https://api.pdf.co/v1/pdf/convert/to/html"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("simple", PlainHtml); parameters.Add("columns", ColumnLayout); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated HTML file string resultFileUrl = json["url"].ToString(); // Download HTML file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated HTML file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination HTML file name final static Path DestinationFile = Paths.get(".\\result.html"); // Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design. final static boolean PlainHtml = false; // Set to `true` if your document has the column layout like a newspaper. final static boolean ColumnLayout = false; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO HTML PdfToHtml(webClient, API_KEY, DestinationFile, Password, Pages, PlainHtml, ColumnLayout, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToHtml(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, boolean plainHtml, boolean columnLayout, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To HTML` API call String query = "https://api.pdf.co/v1/pdf/convert/to/html"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"simple\": \"%s\", \"columns\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, plainHtml, columnLayout, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated HTML file String resultFileUrl = json.get("url").getAsString(); // Download HTML file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated HTML file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To HTML Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function PdfToHtml($apiKey, $uploadedFileUrl, $pages, $plainHtml, $columnLayout) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/html"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; if($plainHtml){ $parameters["simple"] = $plainHtml; } if($columnLayout){ $parameters["columns"] = $columnLayout; } // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# PDF to JPG Source: https://developer.pdf.co/api/pdf-to-image/jpg PDF to high quality JPEG image conversion. High quality rendering. Also works great for thumbnail generation and previews. **Try it live:** [PDF to JPG → API Tester](/api-tester/pdf-to-image/jpg) — send a real request from your browser. ## `POST /v1/pdf/convert/to/jpg` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. | |     `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not | |     `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not | |     `RenderCurveVectorObjects` | boolean | *No* | `true` | Render curve vector objects or not | |     `TextSmoothingMode` | string | *No* | - | Controls text smoothing mode. Available options: `HighSpeed`, `HighQuality`. | |     `VectorSmoothingMode` | string | *No* | - | Controls vector smoothing mode. Available options: `HighSpeed`, `HighQuality`. | |     `ImageInterpolationMode` | string | *No* | - | Controls image interpolation mode. Available options: `HighSpeed`, `HighQuality`. | |     `JPEGQuality` | integer | *No* | 80 | Range from `0` (lowest) to `100` (highest), default is `80`. See [profiles.JPEGQuality](#profiles-jpegquality) | |     `TIFFCompression` | string | *No* | - | Controls TIFF compression. Available options: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE`. | |     `RotateFlipType` | string | *No* | - | Controls rotation and flip type. Available options: `RotateNoneFlipNone`, `Rotate90FlipNone`, `Rotate180FlipNone`, `Rotate270FlipNone`, `RotateNoneFlipX`, `Rotate90FlipX`, `Rotate180FlipX`, `Rotate270FlipX`, `RotateNoneFlipY`, `Rotate90FlipY`, `Rotate180FlipY`, `Rotate270FlipY`, `RotateNoneFlipXY`, `Rotate90FlipXY`, `Rotate180FlipXY`, `Rotate270FlipXY`. | |     `ImageBitsPerPixel` | string | *No* | - | Controls image bits per pixel. Available options: `BPP1`, `BPP8`, `BPP24`, `BPP32`. | |     `OneBitConversionAlgorithm` | string | *No* | - | Controls one-bit conversion algorithm. Available options: `BayerOrderedDithering`, `OtsuThreshold`. | |     `FontHintingMode` | string | *No* | - | Controls font hinting mode. Available options: `Default`, `Stronger`. | |     `ResolutionOverride` | float | *No* | - | Overrides the default resolution. Specified in DPI. | |     `NightMode` | boolean | *No* | `false` | Enables night mode rendering. | |     `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Disable Text Layer We can turn off the text layer for our render as follows: ```json theme={null} { "profiles": "{ 'RenderTextObjects': false }" } ``` ### Set Image Resolution By default the screen resolution is 120 DPI. To change the rendering resolution, please use: ```json theme={null} { "profiles": "{ 'RenderingResolution': 300 }" } ``` ### `JPEGQuality` To set image quality (from `0` (lowest) to `100` (highest), default is `80`) please use: ``` { "profiles": "{ 'JPEGQuality': 75 }" } ``` ### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more. Profiles section for more. ```json theme={null} { "profiles": "{ 'RenderTextObjects': false, 'RenderVectorObjects': true, 'RenderImageObjects': true }" } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png", "https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png" ], "pageCount": 2, "error": false, "status": 200, "name": "sample.png", "duration": 1121, "remainingCredits": 98722216, "credits": 30 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO IMAGE convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg"); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } // `imageType` should correspond to the API we want to use, // i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff // we just take the last part of the path, the file extension function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) { // Prepare URL for PDF to Image API call var queryPath = `/v1/pdf/convert/to/${imageType}`; // JSON payload for api request var jsonPayload = JSON.stringify({ password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download generated image files var page = 1; data.urls.forEach((url) => { var localFileName = `./page${page}.${imageType}`; var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated image file saved as "${localFileName}" file.`); }); }); page++; }, this); } else { // Service reported error console.log("convertPDFToImage(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPDFToImage(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPDFToImage(uploadedFileUrl, "jpg") def convertPDFToImage(uploadedFileUrl, imageType): """Converts PDF To Image using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-image/jpg parameters = {} parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for PDF To Image API request url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Download generated JPG files part = 1 for resultFileUrl in json["urls"]: # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.{imageType}" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; const string imageType = "jpg"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have the direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); // Get URL of uploaded file to use with later API calls string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. CONVERT UPLOADED PDF FILE TO IMAGE // Prepare URL for PDF To Image API call string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated JPEG files int page = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\page{0}.{1}", page, imageType); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); page++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO IMAGE PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg"); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException { // Prepare URL for PDF To Image API call String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated JPEG files JsonArray urls = json.get("urls").getAsJsonArray(); int page = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\page%s.%s", page, imageType); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Downloaded \"%s\".", localFileName)); page++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Image Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function RenderPDF($apiKey, $fileUrl, $outputFormat, $pages) { $formats = array(0 => "png", 1 => "jpg", 2 => "tiff"); $format = $formats[$outputFormat]; // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/" . $format; // Prepare requests params $parameters = array(); $parameters["url"] = $fileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { if ($format == "tiff") { // The TIFF format is multi-page, so the result is always the single file $resultFileUrl = $json["url"]; echo "

" . $resultFileUrl . "

"; } else { // JPEG and PNG formats are single-page, so the results are multiple $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# PDF to PNG Source: https://developer.pdf.co/api/pdf-to-image/png PDF to high quality PNG image conversion. High quality rendering. Also works great for thumbnail generation and previews. **Try it live:** [PDF to PNG → API Tester](/api-tester/pdf-to-image/png) — send a real request from your browser. ## `POST /v1/pdf/convert/to/png` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ------------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. | |     `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not | |     `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not | |     `RenderCurveVectorObjects` | boolean | *No* | `true` | Render curve vector objects or not | |     `TextSmoothingMode` | string | *No* | - | Controls text smoothing mode. Available options: `HighSpeed`, `HighQuality`. | |     `VectorSmoothingMode` | string | *No* | - | Controls vector smoothing mode. Available options: `HighSpeed`, `HighQuality`. | |     `ImageInterpolationMode` | string | *No* | - | Controls image interpolation mode. Available options: `HighSpeed`, `HighQuality`. | |     `TIFFCompression` | string | *No* | - | Controls TIFF compression. Available options: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE`. | |     `RotateFlipType` | string | *No* | - | Controls rotation and flip type. Available options: `RotateNoneFlipNone`, `Rotate90FlipN one`, `Rotate180FlipNone`, `Rotate270FlipNone`, `RotateNoneFlipX`, `Rotate90FlipX`, `Rotate180FlipX`, `Rotate270FlipX`, `RotateNoneFlipY`, `Rotate90FlipY`, `Rotate180FlipY`, `Rotate270FlipY`, `RotateNoneFlipXY`, `Rotate90FlipXY`, `Rotate180FlipXY`, `Rotate270FlipXY`. | |     `ImageBitsPerPixel` | string | *No* | - | Controls image bits per pixel. Available options: `BPP1`, `BPP8`, `BPP24`, `BPP32`. | |     `OneBitConversionAlgorithm` | string | *No* | - | Controls one-bit conversion algorithm. Available options: `BayerOrderedDithering`, `FloydSteinbergDithering`, `Threshold`, `OrderedDithering`. | |     `FontHintingMode` | string | *No* | - | Controls font hinting mode. Available options: `Default`, `Stronger`. | |     `ResolutionOverride` | float | *No* | - | Overrides the default resolution. Specified in DPI. | |     `NightMode` | boolean | *No* | `false` | Enables night mode rendering. | |     `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Disable Text Layer We can turn off the text layer for our render as follows: ```json theme={null} { "profiles": "{ 'RenderTextObjects': false }" } ``` ### Set Image Resolution By default the screen resolution is 120 DPI. To change the rendering resolution, please use: ```json theme={null} { "profiles": "{ 'RenderingResolution': 300 }" } ``` ### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more. Profiles section for more. ``` "profiles": { 'RenderTextObjects': false, 'RenderVectorObjects': true, 'RenderImageObjects': true } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png", "https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png" ], "pageCount": 2, "error": false, "status": 200, "name": "sample.png", "duration": 1121, "remainingCredits": 98722216, "credits": 30 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO IMAGE convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg"); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } // `imageType` should correspond to the API we want to use, // i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff // we just take the last part of the path, the file extension function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) { // Prepare URL for PDF to Image API call var queryPath = `/v1/pdf/convert/to/${imageType}`; // JSON payload for api request var jsonPayload = JSON.stringify({ password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download generated image files var page = 1; data.urls.forEach((url) => { var localFileName = `./page${page}.${imageType}`; var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated image file saved as "${localFileName}" file.`); }); }); page++; }, this); } else { // Service reported error console.log("convertPDFToImage(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPDFToImage(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPDFToImage(uploadedFileUrl, "jpg") def convertPDFToImage(uploadedFileUrl, imageType): """Converts PDF To Image using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-image/png parameters = {} parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for PDF To Image API request url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Download generated JPG files part = 1 for resultFileUrl in json["urls"]: # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.{imageType}" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; const string imageType = "jpg"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have the direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); // Get URL of uploaded file to use with later API calls string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. CONVERT UPLOADED PDF FILE TO IMAGE // Prepare URL for PDF To Image API call string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated JPEG files int page = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\page{0}.{1}", page, imageType); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); page++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO IMAGE PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg"); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException { // Prepare URL for PDF To Image API call String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated JPEG files JsonArray urls = json.get("urls").getAsJsonArray(); int page = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\page%s.%s", page, imageType); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Downloaded \"%s\".", localFileName)); page++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Image Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function RenderPDF($apiKey, $fileUrl, $outputFormat, $pages) { $formats = array(0 => "png", 1 => "jpg", 2 => "tiff"); $format = $formats[$outputFormat]; // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/" . $format; // Prepare requests params $parameters = array(); $parameters["url"] = $fileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { if ($format == "tiff") { // The TIFF format is multi-page, so the result is always the single file $resultFileUrl = $json["url"]; echo "

" . $resultFileUrl . "

"; } else { // JPEG and PNG formats are single-page, so the results are multiple $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to TIFF Source: https://developer.pdf.co/api/pdf-to-image/tiff PDF to high quality TIFF image conversion. High quality rendering. Also works great for thumbnail generation and previews. **Try it live:** [PDF to TIFF → API Tester](/api-tester/pdf-to-image/tiff) — send a real request from your browser. ## `POST /v1/pdf/convert/to/tiff` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `TIFFCompression` | string | *No* | `LZW` | See [profiles.TIFFCompression](#profiles-tiffcompression) | |     `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. | |     `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Disable Text Layer We can turn off the text layer for our render as follows: ```json theme={null} { "profiles": "{ 'RenderTextObjects': false }" } ``` ### Set Image Resolution By default the screen resolution is 120 DPI. To change the rendering resolution, please use: ```json theme={null} { "profiles": "{ 'RenderingResolution': 300 }" } ``` ### `profiles.TIFFCompression` TIFF has a variety of options as follows: ``` { "profiles": "{ 'RenderTextObjects': true, // Valid values: true, false 'RenderVectorObjects': true, // Valid values: true, false 'RenderImageObjects': true, // Valid values: true, false 'TIFFCompression': 'LZW', // Valid values: 'None', 'LZW', 'CCITT3', 'CCITT4', 'RLE' }" } ``` ### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more. Profiles section for more. ``` "profiles": { 'RenderTextObjects': false, 'RenderVectorObjects': true, 'RenderImageObjects': true } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png", "https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png" ], "pageCount": 2, "error": false, "status": 200, "name": "sample.png", "duration": 1121, "remainingCredits": 98722216, "credits": 30 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO IMAGE convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg"); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } // `imageType` should correspond to the API we want to use, // i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff // we just take the last part of the path, the file extension function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) { // Prepare URL for PDF to Image API call var queryPath = `/v1/pdf/convert/to/${imageType}`; // JSON payload for api request var jsonPayload = JSON.stringify({ password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download generated image files var page = 1; data.urls.forEach((url) => { var localFileName = `./page${page}.${imageType}`; var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated image file saved as "${localFileName}" file.`); }); }); page++; }, this); } else { // Service reported error console.log("convertPDFToImage(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPDFToImage(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPDFToImage(uploadedFileUrl, "jpg") def convertPDFToImage(uploadedFileUrl, imageType): """Converts PDF To Image using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-image/tiff parameters = {} parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for PDF To Image API request url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Download generated JPG files part = 1 for resultFileUrl in json["urls"]: # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.{imageType}" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; const string imageType = "jpg"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have the direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); // Get URL of uploaded file to use with later API calls string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. CONVERT UPLOADED PDF FILE TO IMAGE // Prepare URL for PDF To Image API call string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated JPEG files int page = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\page{0}.{1}", page, imageType); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); page++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO IMAGE PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg"); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException { // Prepare URL for PDF To Image API call String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated JPEG files JsonArray urls = json.get("urls").getAsJsonArray(); int page = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\page%s.%s", page, imageType); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Downloaded \"%s\".", localFileName)); page++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Image Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function RenderPDF($apiKey, $fileUrl, $outputFormat, $pages) { $formats = array(0 => "png", 1 => "jpg", 2 => "tiff"); $format = $formats[$outputFormat]; // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/" . $format; // Prepare requests params $parameters = array(); $parameters["url"] = $fileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { if ($format == "tiff") { // The TIFF format is multi-page, so the result is always the single file $resultFileUrl = $json["url"]; echo "

" . $resultFileUrl . "

"; } else { // JPEG and PNG formats are single-page, so the results are multiple $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to WEBP Source: https://developer.pdf.co/api/pdf-to-image/webp PDF to high quality WEBP image conversion. High quality rendering. Also works great for thumbnail generation and previews. **Try it live:** [PDF to WEBP → API Tester](/api-tester/pdf-to-image/webp) — send a real request from your browser. ## `POST /v1/pdf/convert/to/webp` `WEBP` is an image format invented by Google and is supported by Google Chrome and other modern browsers. It provides good quality with smaller file sizes. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` | |     `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. | |     `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. | |     `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not | |     `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not | |     `WEBPQuality` | integer | *No* | 75 | See [profiles.WEBPQuality](#profiles-webpquality) | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Disable Text Layer We can turn off the text layer for our render as follows: ```json theme={null} { "profiles": "{ 'RenderTextObjects': false }" } ``` ### Set Image Resolution By default the screen resolution is 120 DPI. To change the rendering resolution, please use: ```json theme={null} { "profiles": "{ 'RenderingResolution': 300 }" } ``` ### `profiles.WEBPQuality` To control the quality and encoding speed use the following: ``` { "profiles": "{ 'WEBPQuality': 75 }" } ``` ### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more. Profiles section for more. ``` "profiles": { 'RenderTextObjects': false, 'RenderVectorObjects': true, 'RenderImageObjects': true } ``` ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ | | `urls` | array\[string] | List of URLs to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "urls": [ "https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png", "https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png" ], "pageCount": 2, "error": false, "status": 200, "name": "sample.png", "duration": 1121, "remainingCredits": 98722216, "credits": 30 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf", "inline": true, "pages": "0-", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO IMAGE convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg"); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } // `imageType` should correspond to the API we want to use, // i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff // we just take the last part of the path, the file extension function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) { // Prepare URL for PDF to Image API call var queryPath = `/v1/pdf/convert/to/${imageType}`; // JSON payload for api request var jsonPayload = JSON.stringify({ password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": API_KEY, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download generated image files var page = 1; data.urls.forEach((url) => { var localFileName = `./page${page}.${imageType}`; var file = fs.createWriteStream(localFileName); https.get(url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated image file saved as "${localFileName}" file.`); }); }); page++; }, this); } else { // Service reported error console.log("convertPDFToImage(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPDFToImage(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPDFToImage(uploadedFileUrl, "jpg") def convertPDFToImage(uploadedFileUrl, imageType): """Converts PDF To Image using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-image/webp parameters = {} parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for PDF To Image API request url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Download generated JPG files part = 1 for resultFileUrl in json["urls"]: # Download Result File r = requests.get(resultFileUrl, stream=True) localFileUrl = f"Page{part}.{imageType}" if r.status_code == 200: with open(localFileUrl, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{localFileUrl}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") part = part + 1 else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; const string imageType = "jpg"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have the direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); // Get URL of uploaded file to use with later API calls string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream // 3. CONVERT UPLOADED PDF FILE TO IMAGE // Prepare URL for PDF To Image API call string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType); // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Download generated JPEG files int page = 1; foreach (JToken token in json["urls"]) { string resultFileUrl = token.ToString(); string localFileName = String.Format(@".\page{0}.{1}", page, imageType); webClient.DownloadFile(resultFileUrl, localFileName); Console.WriteLine("Downloaded \"{0}\".", localFileName); page++; } } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonArray; import com.google.gson.JsonElement; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO IMAGE PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg"); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException { // Prepare URL for PDF To Image API call String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Download generated JPEG files JsonArray urls = json.get("urls").getAsJsonArray(); int page = 1; for (JsonElement element: urls) { String resultFileUrl = element.getAsString(); String localFileName = String.format(".\\page%s.%s", page, imageType); downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile()); System.out.println(String.format("Downloaded \"%s\".", localFileName)); page++; } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Image Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function RenderPDF($apiKey, $fileUrl, $outputFormat, $pages) { $formats = array(0 => "png", 1 => "jpg", 2 => "tiff"); $format = $formats[$outputFormat]; // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/" . $format; // Prepare requests params $parameters = array(); $parameters["url"] = $fileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { if ($format == "tiff") { // The TIFF format is multi-page, so the result is always the single file $resultFileUrl = $json["url"]; echo "

" . $resultFileUrl . "

"; } else { // JPEG and PNG formats are single-page, so the results are multiple $resultFiles = $json["urls"]; foreach ($resultFiles as &$resultFileUrl) echo "

" . $resultFileUrl . "

"; } } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to JSON Source: https://developer.pdf.co/api/pdf-to-json/basic Convert PDF and scanned images into JSON representation with text, fonts, images, vectors, and formatting preserved. **Try it live:** [PDF to JSON → API Tester](/api-tester/pdf-to-json/basic) — send a real request from your browser. ## `POST /v1/pdf/convert/to/json2` This endpoint can also be used with a specified [profile to extract image data from a PDF into your JSON output](/api/profiles/#api-profiles-save-images). You can extract hyperlinks from your PDF by using a [profile to extract only link objects](/api/profiles#extract-hyperlinks). ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the JSON file. ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | | | ------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | - | - | | `body` | object | *No* | - | - | |     `properties` | object | *No* | - | - | |         `document` | object | *No* | - | - | |             `properties` | object | *No* | - | - | |                 `pageCount` | string | Total number of pages in the document. | | | |                 `pageCountWithOCRPerformed` | string | Total number of pages in the document with OCR performed. | | | |                 `page` | object | Page details. | | | | `pageCount` | integer | Number of pages in the PDF document. | | | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | | | `name` | string | Name of the output file | | | | `credits` | integer | Number of credits consumed by the request | | | | `remainingCredits` | integer | Number of credits remaining in the account | | | | `duration` | integer | Time taken for the operation in milliseconds | | | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf", "inline": true, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "url": "https://example.com/file1.pdf" } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/json2' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf", "inline": true, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination JSON file name const DestinationFile = "./result.json"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO JSON convertPdfToJson(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToJson(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To JSON` API endpoint var queryPath = `/v1/pdf/convert/to/json`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download JSON file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated JSON file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToJson(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToJson(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination JSON file name DestinationFile = ".\\result.json" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToJson(uploadedFileUrl, DestinationFile) def convertPdfToJson(uploadedFileUrl, destinationFile): """Converts PDF To Json using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-json/basic parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Json' API request url = "{}/pdf/convert/to/json".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination JSON file name const string DestinationFile = @".\result.json"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO JSON // URL for `PDF To JSON` API call var url = "https://api.pdf.co/v1/pdf/convert/to/json"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated JSON file string resultFileUrl = json["url"].ToString(); // Download JSON file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination JSON file name final static Path DestinationFile = Paths.get(".\\result.json"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO JSON PdfToJson(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToJson(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To JSON` API call String query = "https://api.pdf.co/v1/pdf/convert/to/json"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated JSON file String resultFileUrl = json.get("url").getAsString(); // Download JSON file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To JSON Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractJSON($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/json"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ``` # PDF to JSON with AI Source: https://developer.pdf.co/api/pdf-to-json/with-ai Convert PDF and scanned images into JSON using AI. **Try it live:** [PDF to JSON with AI → API Tester](/api-tester/pdf-to-json/with-ai) — send a real request from your browser. ## `POST /v1/pdf/convert/to/json-meta` This endpoint can also be used with a specified [profile to extract image data from a PDF into your JSON output](/api/profiles/#api-profiles-save-images). What is the difference between `/pdf/convert/to/json-meta` and `/pdf/convert/to/json2`? `/pdf/convert/to/json-meta` uses AI to detect meta styles for text objects, such as: * paragraph style (from `h1` .. `h7` to `p` and `small`). * meta `type` of the text object (`text`, `datetime`, `integer`, `decimal`, `currency` etc.). * meta `subType` of the text object (`companyName`, `personName` and other AI-based meta types). * `/json-meta` consumes more credits because it runs with AI. * `/json-meta` is also a bit slower due to the AI process. `Async` mode is recommended for this endpoint. Convert PDF and scanned images into JSON using AI. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file. ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | ------------------ | ------- | -------------------------------------- | | `body` | object | Response body. | | `pageCount` | integer | Total number of pages in the document. | | `error` | boolean | Indicates if an error occurred. | | `status` | integer | Status code. | | `name` | string | Name of the job. | | `credits` | integer | Total number of credits used. | | `remainingCredits` | integer | Remaining number of credits. | | `duration` | integer | Duration of the job. | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf", "inline": true, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": { "document": { "pageCount": "1", "pageCountWithOCRPerformed": "0", "page": { "index": "0", "width": "595.320007324219", "height": "841.919982910156", "OCRWasPerformed": "False", "row": [ { "column": [ { "text": { "fontName": "Arial", "fontSize": "24.0", "fontStyle": "Bold", "color": "#538DD3", "x": "36.00", "y": "34.44", "width": "242.81", "height": "24.00", "text": "Your Company Name" } }, { "text": "" }, { "text": "" }, { "text": "" } ] }, { "column": [ { "text": "" }, { "text": "" }, { "text": { "fontName": "Arial", "fontSize": "11.0", "fontStyle": "Bold", "x": "389.11", "y": "425.83", "width": "36.75", "height": "11.04", "text": "TOTAL" } }, { "text": { "fontName": "Arial", "fontSize": "11.0", "fontStyle": "Bold", "x": "525.82", "y": "425.83", "width": "33.62", "height": "11.04", "text": "200.00" } } ] } ] } } }, "pageCount": 1, "error": false, "status": 200, "name": "sample.json", "remainingCredits": 99227903, "credits": 28 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/json-meta' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf", "inline": true, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination JSON file name const DestinationFile = "./result.json"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO JSON convertPdfToJson(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToJson(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To JSON` API endpoint var queryPath = `/v1/pdf/convert/to/json`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download JSON file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated JSON file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToJson(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToJson(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination JSON file name DestinationFile = ".\\result.json" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToJson(uploadedFileUrl, DestinationFile) def convertPdfToJson(uploadedFileUrl, destinationFile): """Converts PDF To Json using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-json/with-ai parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Json' API request url = "{}/pdf/convert/to/json".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination JSON file name const string DestinationFile = @".\result.json"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO JSON // URL for `PDF To JSON` API call var url = "https://api.pdf.co/v1/pdf/convert/to/json"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated JSON file string resultFileUrl = json["url"].ToString(); // Download JSON file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination JSON file name final static Path DestinationFile = Paths.get(".\\result.json"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO JSON PdfToJson(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToJson(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To JSON` API call String query = "https://api.pdf.co/v1/pdf/convert/to/json"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated JSON file String resultFileUrl = json.get("url").getAsString(); // Download JSON file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To JSON Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractJSON($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/json"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# PDF to Text Source: https://developer.pdf.co/api/pdf-to-text/basic Convert PDF and scanned images to text with layout preserved. This method uses OCR and reporoduces layout. **Try it live:** [PDF to Text → API Tester](/api-tester/pdf-to-text/basic) — send a real request from your browser. ## `POST /v1/pdf/convert/to/text` **Auto classification Of incoming documents**: Use the [Document Classifier](/api/document-classifier) endpoint to automatically sort/detect the class of the document based on keywords-based rules. For example, you can define rules to find which vendor provided the document to find which template to apply accordingly. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "inline": true, "async": false }' ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": " Your Company Name \r\n Your Address \r\n City, State Zip \r\n Invoice No. 123456 \r\n Invoice Date 01/01/2016 \r\n Client Name \r\n Address \r\n City, State Zip \r\n\r\n Notes \r\n\r\n\r\n Item Quantity Price Total \r\n Item 1 1 40.00 40.00 \r\n Item 2 2 30.00 60.00 \r\n Item 3 3 20.00 60.00 \r\n Item 4 4 10.00 40.00 \r\n TOTAL 200.00\r\n", "pageCount": 1, "error": false, "status": 200, "name": "sample.txt", "remainingCredits": 99032333, "credits": 21 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "inline": true, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination TXT file name const DestinationFile = "./result.txt"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO TEXT convertPdfToText(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToText(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To Text` API endpoint var queryPath = `/v1/pdf/convert/to/text`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download TXT file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated TXT file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToText(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToText(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination TXT file name DestinationFile = ".\\result.txt" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToText(uploadedFileUrl, DestinationFile) def convertPdfToText(uploadedFileUrl, destinationFile): """Converts PDF To Text using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-text/basic parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Text' API request url = "{}/pdf/convert/to/text".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination TXT file name const string DestinationFile = @".\result.txt"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO TXT // URL for `PDF To TXT` API call var url = "https://api.pdf.co/v1/pdf/convert/to/text"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated TXT file string resultFileUrl = json["url"].ToString(); // Download TXT file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated TXT file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination TXT file name final static Path DestinationFile = Paths.get(".\\result.txt"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO TXT PdfToText(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToText(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To TXT` API call String query = "https://api.pdf.co/v1/pdf/convert/to/text"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated TXT file String resultFileUrl = json.get("url").getAsString(); // Download TXT file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated TXT file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Text Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractText($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/text"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# PDF to Text (Simple) Source: https://developer.pdf.co/api/pdf-to-text/simple Extract plain text from PDF documents using a fast, low-credit method without OCR, layout analysis, or profile-based fine-tuning. **Try it live:** [PDF to Text (Simple) → API Tester](/api-tester/pdf-to-text/simple) — send a real request from your browser. ## `POST /v1/pdf/convert/to/text-simple` **Auto classification Of incoming documents**: Use the [Document Classifier](/api/document-classifier) endpoint to automatically sort/detect the class of the document based on keywords-based rules. For example, you can define rules to find which vendor provided the document to find which template to apply accordingly. ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string / integer | Status of the API response. Returns the string `success` on success. On failure it is either a numeric code such as `400` for endpoint-level errors, or the string `error` for authentication and routing failures, and the numeric code is also returned in the `errorCode` field. For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "inline": true, "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": "Your Company Name \r\nYour Address \r\nCity, State Zip \r\nInvoice No. 123456 \r\nInvoice Date 01/01/2016 \r\nClient Name \r\nAddress \r\nCity, State Zip \r\nNotes \r\nItem Quantity Price Total \r\nItem 1 1 40.00 40.00 \r\nItem 2 2 30.00 60.00 \r\nItem 3 3 20.00 60.00 \r\nItem 4 4 10.00 40.00 \r\nTOTAL 200.00 \r\n", "pageCount": 1, "error": false, "status": "success", "name": "sample.txt", "remainingCredits": 99885491, "credits": 4 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text-simple' \ --header 'Content-Type: application/json' \ --header 'x-api-key: *******************' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "inline": true, "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination TXT file name const DestinationFile = "./result.txt"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO TEXT convertPdfToText(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(localFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(localFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToText(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To Text (Simple)` API endpoint var queryPath = `/v1/pdf/convert/to/text-simple`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download TXT file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated TXT file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToText(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToText(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination TXT file name DestinationFile = ".\\result.txt" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToText(uploadedFileUrl, DestinationFile) def convertPdfToText(uploadedFileUrl, destinationFile): """Converts PDF To Text using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-text/simple parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To Text (Simple)' API request url = "{}/pdf/convert/to/text-simple".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination TXT file name const string DestinationFile = @".\result.txt"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO TXT // URL for `PDF To Text (Simple)` API call var url = "https://api.pdf.co/v1/pdf/convert/to/text-simple"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated TXT file string resultFileUrl = json["url"].ToString(); // Download TXT file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated TXT file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination TXT file name final static Path DestinationFile = Paths.get(".\\result.txt"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO TXT PdfToText(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToText(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To Text (Simple)` API call String query = "https://api.pdf.co/v1/pdf/convert/to/text-simple"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated TXT file String resultFileUrl = json.get("url").getAsString(); // Download TXT file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated TXT file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To Text Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractText($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/text-simple"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# PDF to XML Source: https://developer.pdf.co/api/pdf-to-xml Convert PDF to XML with information about text value, tables, fonts, images, objects positions. **Try it live:** [PDF to XML → API Tester](/api-tester/pdf-to-xml) — send a real request from your browser. ## `POST /v1/pdf/convert/to/xml` ## Attributes Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` | Attribute | Type | Required | Default | Description | | -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) | | `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. | | `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. | | `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. | | `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | | `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | | `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | | `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | | `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | | `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). | | `password` | string | *No* | - | Password for the PDF file. | | `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) | | `name` | string | *No* | - | File name for the generated output, the input must be in string format. | | `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | | `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. | |     `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). | |     `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. | |     `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. | |     `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). | |     `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | |     `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | |     `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | |     `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |         `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. | |         `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. | |     `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. | |     `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | |     `SaveImages` | string | *No* | `None` | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). | |     `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | |     `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | |     `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. | |     `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. | |     `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | |     `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. | You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file. ### `OCRImagePreprocessingFilters` To set image preprocessing filters, please use: ```json theme={null} { "profiles": "{ "ExtractShadowLikeText": false, "OCRMode": "Auto", "OCRImagePreprocessingFilters.AddGrayscale()": [], "OCRImagePreprocessingFilters.AddGammaCorrection()": [ 1.4 ] }" } ``` ### Line Grouping Options * `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row. * `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines. * `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content. ## Query parameters *No query parameters accepted.* ## Responses | Parameter | Type | Description | | --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | | `url` | string | Direct URL to the final PDF file stored in S3. | | `outputLinkValidTill` | string | Timestamp indicating when the output link will expire | | `pageCount` | integer | Number of pages in the PDF document. | | `error` | boolean | Indicates whether an error occurred (`false` means success) | | `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | `name` | string | Name of the output file | | `credits` | integer | Number of credits consumed by the request | | `remainingCredits` | integer | Number of credits remaining in the account | | `duration` | integer | Time taken for the operation in milliseconds | ## `Example` Payload To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size). ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-xml/sample.pdf", "async": false } ``` ## `Example` Response To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. ```json theme={null} { "body": "\r\n\r\n \r\n \r\n \r\n Your Company Name\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Your Address\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n City, State Zip\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Invoice No. 123456\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Invoice Date 01/01/2016\r\n \r\n \r\n \r\n \r\n Client Name\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Address\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n City, State Zip\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Notes\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Item\r\n \r\n \r\n Quantity\r\n \r\n \r\n Price\r\n \r\n \r\n Total\r\n \r\n \r\n \r\n \r\n Item 1\r\n \r\n \r\n 1\r\n \r\n \r\n 40.00\r\n \r\n \r\n 40.00\r\n \r\n \r\n \r\n \r\n Item 2\r\n \r\n \r\n 2\r\n \r\n \r\n 30.00\r\n \r\n \r\n 60.00\r\n \r\n \r\n \r\n \r\n Item 3\r\n \r\n \r\n 3\r\n \r\n \r\n 20.00\r\n \r\n \r\n 60.00\r\n \r\n \r\n \r\n \r\n Item 4\r\n \r\n \r\n 4\r\n \r\n \r\n 10.00\r\n \r\n \r\n 40.00\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n TOTAL\r\n \r\n \r\n 200.00\r\n \r\n \r\n \r\n", "pageCount": 1, "error": false, "status": 200, "name": "sample.xml", "remainingCredits": 60563 } ``` **Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically. ## Code Samples ```bash theme={null} curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/xml' \ --header 'x-api-key: *******************' \ --header 'Content-Type: application/json' \ --data-raw '{ "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-xml/sample.pdf", "async": false }' ``` ```javascript theme={null} var https = require("https"); var path = require("path"); var fs = require("fs"); // `request` module is required for file upload. // Use "npm install request" command to install. var request = require("request"); // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const API_KEY = "***********************************"; // Source PDF file const SourceFile = "./sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const Pages = ""; // PDF document password. Leave empty for unprotected documents. const Password = ""; // Destination XML file name const DestinationFile = "./result.xml"; // 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. getPresignedUrl(API_KEY, SourceFile) .then(([uploadUrl, uploadedFileUrl]) => { // 2. UPLOAD THE FILE TO CLOUD. uploadFile(API_KEY, SourceFile, uploadUrl) .then(() => { // 3. CONVERT UPLOADED PDF FILE TO XML convertPdfToXml(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile); }) .catch(e => { console.log(e); }); }) .catch(e => { console.log(e); }); function getPresignedUrl(apiKey, localFile) { return new Promise(resolve => { // Prepare request to `Get Presigned URL` API endpoint let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`; let reqOptions = { host: "api.pdf.co", path: encodeURI(queryPath), headers: { "x-api-key": API_KEY } }; // Send request https.get(reqOptions, (response) => { response.on("data", (d) => { let data = JSON.parse(d); if (data.error == false) { // Return presigned url we received resolve([data.presignedUrl, data.url]); } else { // Service reported error console.log("getPresignedUrl(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("getPresignedUrl(): " + e); }); }); } function uploadFile(apiKey, localFile, uploadUrl) { return new Promise(resolve => { fs.readFile(SourceFile, (err, data) => { request({ method: "PUT", url: uploadUrl, body: data, headers: { "Content-Type": "application/octet-stream" } }, (err, res, body) => { if (!err) { resolve(); } else { console.log("uploadFile() request error: " + e); } }); }); }); } function convertPdfToXml(apiKey, uploadedFileUrl, password, pages, destinationFile) { // Prepare request to `PDF To XML` API endpoint var queryPath = `/v1/pdf/convert/to/xml`; // JSON payload for api request var jsonPayload = JSON.stringify({ name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl }); var reqOptions = { host: "api.pdf.co", method: "POST", path: queryPath, headers: { "x-api-key": apiKey, "Content-Type": "application/json", "Content-Length": Buffer.byteLength(jsonPayload, 'utf8') } }; // Send request var postRequest = https.request(reqOptions, (response) => { response.on("data", (d) => { response.setEncoding("utf8"); // Parse JSON response let data = JSON.parse(d); if (data.error == false) { // Download XML file var file = fs.createWriteStream(destinationFile); https.get(data.url, (response2) => { response2.pipe(file) .on("close", () => { console.log(`Generated XML file saved as "${destinationFile}" file.`); }); }); } else { // Service reported error console.log("convertPdfToXml(): " + data.message); } }); }) .on("error", (e) => { // Request error console.log("convertPdfToXml(): " + e); }); // Write request data postRequest.write(jsonPayload); postRequest.end(); } ``` ```python theme={null} import os import requests # pip install requests # The authentication key (API Key). # Get your own by registering at https://app.pdf.co API_KEY = "******************************************" # Base URL for PDF.co Web API requests BASE_URL = "https://api.pdf.co/v1" # Source PDF file SourceFile = ".\\sample.pdf" # Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. Pages = "" # PDF document password. Leave empty for unprotected documents. Password = "" # Destination XML file name DestinationFile = ".\\result.xml" def main(args = None): uploadedFileUrl = uploadFile(SourceFile) if (uploadedFileUrl != None): convertPdfToXml(uploadedFileUrl, DestinationFile) def convertPdfToXml(uploadedFileUrl, destinationFile): """Converts PDF To XML using PDF.co Web API""" # Prepare requests params as JSON # See documentation: https://developer.pdf.co/api/pdf-to-xml parameters = {} parameters["name"] = os.path.basename(destinationFile) parameters["password"] = Password parameters["pages"] = Pages parameters["url"] = uploadedFileUrl # Prepare URL for 'PDF To XML' API request url = "{}/pdf/convert/to/xml".format(BASE_URL) # Execute request and get response as JSON response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # Get URL of result file resultFileUrl = json["url"] # Download result file r = requests.get(resultFileUrl, stream=True) if (r.status_code == 200): with open(destinationFile, 'wb') as file: for chunk in r: file.write(chunk) print(f"Result file saved as \"{destinationFile}\" file.") else: print(f"Request error: {response.status_code} {response.reason}") else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") def uploadFile(fileName): """Uploads file to the cloud""" # 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE. # Prepare URL for 'Get Presigned URL' API request url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format( BASE_URL, os.path.basename(fileName)) # Execute request and get response as JSON response = requests.get(url, headers={ "x-api-key": API_KEY }) if (response.status_code == 200): json = response.json() if json["error"] == False: # URL to use for file upload uploadUrl = json["presignedUrl"] # URL for future reference uploadedFileUrl = json["url"] # 2. UPLOAD FILE TO CLOUD. with open(fileName, 'rb') as file: requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" }) return uploadedFileUrl else: # Show service reported error print(json["message"]) else: print(f"Request error: {response.status_code} {response.reason}") return None if __name__ == '__main__': main() ``` ```csharp theme={null} using System; using System.Collections.Generic; using System.IO; using System.Net; using Newtonsoft.Json; using Newtonsoft.Json.Linq; namespace PDFcoApiExample { class Program { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co const String API_KEY = "***********************************"; // Source PDF file const string SourceFile = @".\sample.pdf"; // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. const string Pages = ""; // PDF document password. Leave empty for unprotected documents. const string Password = ""; // Destination XML file name const string DestinationFile = @".\result.xml"; static void Main(string[] args) { // Create standard .NET web client instance WebClient webClient = new WebClient(); // Set API Key webClient.Headers.Add("x-api-key", API_KEY); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call string query = Uri.EscapeUriString(string.Format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}", Path.GetFileName(SourceFile))); try { // Execute request string response = webClient.DownloadString(query); // Parse JSON response JObject json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL to use for the file upload string uploadUrl = json["presignedUrl"].ToString(); string uploadedFileUrl = json["url"].ToString(); // 2. UPLOAD THE FILE TO CLOUD. webClient.Headers.Add("content-type", "application/octet-stream"); webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream webClient.Headers.Remove("content-type"); // 3. CONVERT UPLOADED PDF FILE TO XML // URL for `PDF To XML` API call var url = "https://api.pdf.co/v1/pdf/convert/to/xml"; // Prepare requests params as JSON Dictionary parameters = new Dictionary(); parameters.Add("name", Path.GetFileName(DestinationFile)); parameters.Add("password", Password); parameters.Add("pages", Pages); parameters.Add("url", uploadedFileUrl); // Convert dictionary of params to JSON string jsonPayload = JsonConvert.SerializeObject(parameters); // Execute POST request with JSON payload response = webClient.UploadString(url, jsonPayload); // Parse JSON response json = JObject.Parse(response); if (json["error"].ToObject() == false) { // Get URL of generated XML file string resultFileUrl = json["url"].ToString(); // Download XML file webClient.DownloadFile(resultFileUrl, DestinationFile); Console.WriteLine("Generated XML file saved as \"{0}\" file.", DestinationFile); } else { Console.WriteLine(json["message"].ToString()); } } else { Console.WriteLine(json["message"].ToString()); } } catch (WebException e) { Console.WriteLine(e.ToString()); } webClient.Dispose(); Console.WriteLine(); Console.WriteLine("Press any key..."); Console.ReadKey(); } } } ``` ```java theme={null} package com.company; import com.google.gson.JsonObject; import com.google.gson.JsonParser; import okhttp3.*; import java.io.*; import java.net.*; import java.nio.file.Path; import java.nio.file.Paths; public class Main { // The authentication key (API Key). // Get your own by registering at https://app.pdf.co final static String API_KEY = "***********************************"; // Source PDF file final static Path SourceFile = Paths.get(".\\sample.pdf"); // Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'. final static String Pages = ""; // PDF document password. Leave empty for unprotected documents. final static String Password = ""; // Destination XML file name final static Path DestinationFile = Paths.get(".\\result.xml"); public static void main(String[] args) throws IOException { // Create HTTP client instance OkHttpClient webClient = new OkHttpClient(); // 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE. // * If you already have a direct file URL, skip to the step 3. // Prepare URL for `Get Presigned URL` API call String query = String.format( "https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s", SourceFile.getFileName()); // Prepare request Request request = new Request.Builder() .url(query) .addHeader("x-api-key", API_KEY) // (!) Set API Key .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL to use for the file upload String uploadUrl = json.get("presignedUrl").getAsString(); // Get URL of uploaded file to use with later API calls String uploadedFileUrl = json.get("url").getAsString(); // 2. UPLOAD THE FILE TO CLOUD. if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile)) { // 3. CONVERT UPLOADED PDF FILE TO XML PdfToXml(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl); } } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static void PdfToXml(OkHttpClient webClient, String apiKey, Path destinationFile, String password, String pages, String uploadedFileUrl) throws IOException { // Prepare URL for `PDF To XML` API call String query = "https://api.pdf.co/v1/pdf/convert/to/xml"; // Make correctly escaped (encoded) URL URL url = null; try { url = new URI(null, query, null).toURL(); } catch (URISyntaxException e) { e.printStackTrace(); } // Create JSON payload String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}", destinationFile.getFileName(), password, pages, uploadedFileUrl); // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", API_KEY) // (!) Set API Key .addHeader("Content-Type", "application/json") .post(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); if (response.code() == 200) { // Parse JSON response JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject(); boolean error = json.get("error").getAsBoolean(); if (!error) { // Get URL of generated XML file String resultFileUrl = json.get("url").getAsString(); // Download XML file downloadFile(webClient, resultFileUrl, destinationFile.toFile()); System.out.printf("Generated XML file saved as \"%s\" file.", destinationFile.toString()); } else { // Display service reported error System.out.println(json.get("message").getAsString()); } } else { // Display request error System.out.println(response.code() + " " + response.message()); } } public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException { // Prepare request body RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile()); // Prepare request Request request = new Request.Builder() .url(url) .addHeader("x-api-key", apiKey) // (!) Set API Key .addHeader("content-type", "application/octet-stream") .put(body) .build(); // Execute request Response response = webClient.newCall(request).execute(); return (response.code() == 200); } public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException { // Prepare request Request request = new Request.Builder() .url(url) .build(); // Execute request Response response = webClient.newCall(request).execute(); byte[] fileBytes = response.body().bytes(); // Save downloaded bytes to file OutputStream output = new FileOutputStream(destinationFile); output.write(fileBytes); output.flush(); output.close(); response.close(); } } ``` ```php theme={null} PDF To XML Extraction Results Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } } else { // Display service reported error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } curl_close($curl); } else { // Display CURL error echo "Error: " . curl_error($curl); } function ExtractXML($apiKey, $uploadedFileUrl, $pages) { // Create URL $url = "https://api.pdf.co/v1/pdf/convert/to/xml"; // Prepare requests params $parameters = array(); $parameters["url"] = $uploadedFileUrl; $parameters["pages"] = $pages; // Create Json payload $data = json_encode($parameters); // Create request $curl = curl_init(); curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json")); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_POST, true); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_POSTFIELDS, $data); // Execute request $result = curl_exec($curl); if (curl_errno($curl) == 0) { $status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE); if ($status_code == 200) { $json = json_decode($result, true); if (!isset($json["error"]) || $json["error"] == false) { $resultFileUrl = $json["url"]; // Display link to the file with conversion results echo "

Conversion Result:

" . $resultFileUrl . "
"; } else { // Display service reported error echo "

Error: " . $json["message"] . "

"; } } else { // Display request error echo "

Status code: " . $status_code . "

"; echo "

" . $result . "

"; } } else { // Display CURL error echo "Error: " . curl_error($curl); } // Cleanup curl_close($curl); } ?> ```
# Postman Source: https://developer.pdf.co/api/postman Import the PDF.co Postman Collection to explore and test every API endpoint. [Postman](https://www.postman.com/) is the collaboration platform for API development, used by 10 million developers and 500,000 companies worldwide. The Postman API Platform simplifies each step of building an API, and streamlines collaboration, so you can create better APIs — faster. ## Getting Started with Postman & PDF.co We have created **PDF.co Collection** for **Postman** that you can import into **Postman** and explore **PDF.co API** and functions right away. Before beginning the **Postman Collection** tutorial, please do the following: * Download and install the [Postman](https://www.postman.com/) app. * Download the [PDF.co Postman Collection](https://pdfco-docs-mintlify.s3.ap-southeast-2.amazonaws.com/PDF.co+API+v.1.00.postman_collection.json). * Import this `.JSON` file into Postman and set up your API key for use with tests. Make sure that you have your API Key ready. # Profiles Source: https://developer.pdf.co/api/profiles This page describes the `profiles` parameter that can be used with your API calls. Profiles are used to to set extra options for common API calls and are sometimes distinct to a particular API. Profiles are embedded with a `JSON` type of notation along with the `profiles` object for your API calls, for example: Please note that the value for the `profiles` field in the code snippets must be enclosed in quotes (`"`), making it a complete string. For example: `{ "profiles": "{'TrimSpaces':true, 'PreserveFormattingOnTextExtraction': true}"}` ## Sample Code ```json theme={null} { "profiles": "{'TrimSpaces':true, 'PreserveFormattingOnTextExtraction': true}" } ``` ``` profiles = '"TrimSpaces": "True", "PreserveFormattingOnTextExtraction": "True" ' ``` ```json theme={null} { "profiles": "'TrimSpaces': 'True' , 'PreserveFormattingOnTextExtraction': 'True'" } ``` ``` String profiles = "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }"; ``` ``` const Profiles = "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }"; ``` ``` $Profiles = '{ "TrimSpaces": "True", "PreserveFormattingOnTextExtraction": "True" }' ``` ```json theme={null} { "profiles": "'TrimSpaces': 'True' , 'PreserveFormattingOnTextExtraction': 'True'" } ``` ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf", "inline": true, "profiles": "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }" } ``` *** ## Generic Profile Options The following `profiles` options are not specific to any one particular endpoint. ### Standard Parameters The `std_params` within the `profiles` parameter enables the definition of regular API parameters in a `JSON` format. This `std_params` feature is designed to simplify the process of passing standard parameters and additional options in the `profiles` parameter for PDF.co API requests. When using [Standard Parameters](#standard-parameters) webhooks can be utilized by setting the `callback` object with the URL of your choice. However, is is simpler to set the `callback` object directly - see [Webhooks & Callbacks](/api/webhooks) for more. When `std_params` are used in the `profiles` parameter, if a parameter is duplicated within both `std_params` and outside profiles, the value specified in `std_params` will overwrite the duplicate value. Therefore if you define a callback object in `std_params` then it will overwrite any value you may have defined via [the basic callback object](/api/webhooks)! #### `std_params` Structure * **Description**: Contains key-value pairs of standard parameters that will be used across PDF.co API requests. * **Type**: `JSON` Object (passed as a string) * **Example**: ```json theme={null} { "profiles": "{'std_params': {'callback': 'webhook_url'}}" } ``` #### Practical Application Using the `std_params` profile, you can define a set of standard parameters and configurations that will be consistently applied across your PDF.co API requests. This approach is particularly beneficial when using automation platforms like [Zapier](/integrations/zapier/), [Make](/integrations/make/), and others, where the number of parameters you can pass directly is limited. #### Complete Request Example Here is a complete example illustrating the use of the `std_params` profile with other parameters: /pdf/convert/to/text ```json theme={null} { "url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf", "inline": true, "profiles": "{'std_params': {'callback': 'webhook_url', 'async': true}, 'ExtractShadowLikeText': false, 'ExtractColumnByColumn': true, 'OCRMode': 'Auto'}}", "TrimSpaces": true, "PreserveFormattingOnTextExtraction": true } ``` ### Output as Base64 If you require your output as `base64` use the following: ```json theme={null} { "profiles": "{ 'outputDataFormat': 'base64' }" } ``` This output data format is supported by endpoints that generate binary files - **PDF** and images. The output is accessible via a generated link and the file under the link is in a base64-encoded text format. ### Converting PDFs There are a variety of `profiles` options which can be set when converting from **PDF** to other documents. These `profiles` control how to extract the information from the source **PDF** file. These options apply to the following endpoints: * /pdf/convert/to/csv * /pdf/convert/to/xml * /pdf/convert/to/json * /pdf/convert/to/json2 * /pdf/convert/to/xls * /pdf/convert/to/xlsx #### Convert Vectors You can choose whether the conversion process should convert vectors or not as follows: ```json theme={null} { "profiles": "{ 'SaveVectors': true }" } ``` #### Save Images This `profiles` parameter includes the `SaveImages` property that extracts individual images in a regular **PDF**. ```json theme={null} { "profiles": "{ 'SaveImages': 'Embed' }" } ``` #### Consider Font Size This `profiles` parameter allows you to seperate header and body text based on font size. ```json theme={null} { "profiles": "{ 'ConsiderFontSizes': true }" } ``` #### Set the Extraction Area Extract text in a specific area by defining the extraction area - set with points in the format `[x, y, width, height]`. ```json theme={null} { "profiles": "{ 'ExtractionArea': [171.0,69.0,249.75,71.25] }" } ``` #### Extract Hyperlinks Extract hyperlinks (URLs) from a PDF document by using the `OutputStructure` and `OutputTransformation` profile options. This returns only the link objects found in the PDF. ```json theme={null} { "profiles": "{ 'OutputStructure': 'OnlyLinks', 'OutputTransformation': '$..text' }" } ``` * `OutputStructure`: Set to `OnlyLinks` to restrict the output to hyperlink elements only. * `OutputTransformation`: Set to `$..text` (a JSONPath expression) to extract just the link text/URL values from the result. #### Extracting Invisible Text When dealing with **PDF** documents, sometimes there may be unwanted invisible text that makes it difficult to extract the desired content accurately. This could be due to various reasons such as the original document being scanned or saved with a low-quality setting. In such cases, it is important to remove the unwanted invisible text to ensure accurate extraction of the desired content. ```json theme={null} { "profiles": "{ 'ExtractInvisibleText': false, 'ExtractShadowLikeText': false, 'OCRMode': 'Auto' }" } ``` ### OCR (Optical Character Recognition) Mode Options The following values can be configured for OCR mode: | OCR Mode | Description | | ------------------------------------------ | ----------------------------------------------------------------------------- | | `Auto` **(default)** | Automatically determines the optimal OCR settings based on the input. | | `AutoRepairFonts` | Automatically repairs fonts in text extracted from images or other documents. | | `TextFromImagesAndFonts` | Extracts text from images and fonts from documents. | | `TextFromImagesAndRepairedFonts` | Extracts text from images and repaired fonts from documents. | | `TextFromImagesAndVectorsAndFonts` | Extracts text, vectors, and fonts from images and documents. | | `TextFromImagesAndVectorsAndRepairedFonts` | Extracts text, vectors, and repaired fonts from images and documents. | | `TextFromImagesAndVectorsOnly` | Extracts text and vectors from images only. | | `TextFromImagesOnly` | Extracts text from images only. | | `TextFromRepairedFontsOnly` | Extracts text from documents with repaired fonts only. | | `TextFromVectorsAndFonts` | Extracts text and fonts from documents with vectors. | | `TextFromVectorsAndRepairedFonts` | Extracts text and repaired fonts from documents with vectors. | | `TextFromVectorsOnly` | Extracts text from documents with vectors only. | ```json theme={null} { "profiles": "{ 'OCRMode': 'TextFromImagesAndVectorsAndRepairedFonts' }" } ``` ### OCR (Optical Character Recognition) Resolution OCR resolution can be set from `72` to `1200` DPI. The default value is `300` DPI. The higher the resolution, the better the OCR results. However, higher resolution also means longer processing times. ```json theme={null} { "profiles": "{ 'OCRResolution': 300 }" } ``` #### Extracting Text from Colored Background If you can’t extract text with a colored background, please add the Grayscale filter to the `profiles` as follows: ```json theme={null} { "profiles": "{ 'OCRImagePreprocessingFilters.AddGrayscale()': [] }" } ``` #### Considering the Font Color on Tables Sometimes the data which OCR must extract from a table might have colored text which is difficult to extract. OCR results can be improved with the following: ```json theme={null} { "profiles": "{ 'LineGroupingMode': 'JoinOrphanedRows', 'ConsiderFontColors': true, 'DetectNewColumnBySpacesRatio': '1.1', 'AutoAlignColumnsToHeader': false, 'OCRImagePreprocessingFilters.AddGammaCorrection()': [ '1.4' ] }" } ``` #### Setting the Rotation Angle Normally OCR detects **PDF** rotation and extracts text properly. But in some cases a **PDF** is constructed in such a way that a page is not rotated and instead text is drawn vertically, OCR does not detect page rotation automatically. In such scenarios we can use following profile setting. ```json theme={null} { "profiles": "{ 'RotationAngle': 2 }" } ``` * `0` no rotation * `1` 90 degrees * `2` 180 degrees * `3` 270 degrees # Response Codes Source: https://developer.pdf.co/api/response-codes Reference list of HTTP status codes and PDF.co-specific error codes returned by the API, with the meaning of each code. The full set of response codes from the **PDF.co** API are as follows: | Code | Description | | ----- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `200` | Success. | | `204` | No content. The server successfully processed the request, and is not returning any content. | | `400` | Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password). | | `401` | Unauthorized. Authentication is required and has failed or has not yet been provided. | | `402` | Not enough credits. | | `403` | Access forbidden for input URL. | | `404` | The requested resource could not be found. | | `408` | The server timed out waiting for the request. | | `414` | The URI provided was too long for the server to process. | | `415` | The request entity has a media type not supported by the server or resource. | | `429` | Too many requests in a given time period. | | `441` | Invalid Password. Password protected document. | | `442` | Input document is damaged or of incorrect type. | | `443` | Permissions. The operation is prohibited by document security settings. You can turn off this check by setting the `profiles` param to `{CheckPermissions: false}`. **Important:** only use this if you are the owner or have legal permission. | | `444` | Profiles parsing error. Please ensure that the configuration is supported. See `/profiles` samples. | | `445` | Timeout error. For large documents, use asynchronous mode (`async=true`) and check status via `/job/check`. For many-page files, use the `pages` parameter. | | `446` | Some files required for conversion are missing. | | `447` | Invalid template. | | `448` | Invalid URL or HTML. Ensure the provided URL is valid and accessible. | | `449` | Invalid index range. Page index is out of range. Use `/pdf/info` to get page count. First page is `0`. | | `450` | Invalid page range specified. | | `452` | Invalid URL. | | `454` | Invalid parameters. | | `455` | Failed to send email. | | `456` | Invalid color. Should be a valid name (e.g., "Red") or hex string (e.g., "#CCBBAA" or "CCBBAA"). | | `457` | SMTP server blocked. | | `466` | Invalid base64 image. | | `490` | Incorrect result data. | | `500` | Something went wrong. Please try again or contact support. | | `501` | Not implemented. The server was acting as a gateway or proxy and received an invalid response. | | `502` | Bad gateway. The server was acting as a gateway or proxy and received an invalid response. | | `503` | Service unavailable. Server is overloaded or under maintenance. | | `504` | Gateway timeout. Server didn’t get a timely response from upstream. | | `505` | HTTP version not supported. | # URL Input and Request Limits Source: https://developer.pdf.co/api/url-input-and-request-limits Supported URL sources, TLS requirements, request size limits, and the cache prefix option for reusable file inputs in PDF.co API calls. ## Supported File Sources The API supports TLS 1.2 and 1.3 for secure connections. Earlier versions such as TLS 1.0 and 1.1 are deprecated and should be avoided. **Tip:** For re-usable files (e.g., PDF templates), use `cache:` before the URL\ (e.g., `cache:https://example.com/file1.pdf`) to reduce repeated downloads and avoid\ errors like `Access Denied or Too Many Requests`. It stores a local copy of the file after the first download, so it does not need to be fetched again from the original server. The **PDF.co API** supports publicly accessible links from any source, including [Google Drive](https://drive.google.com), [Dropbox](https://dropbox.com), and [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). File inputs are accepted only via URLs and not through direct uploads. If your file is stored locally or not publicly accessible, you must upload it using the [File Upload](/api/file-upload) endpoints to get a publicly accessible URL. **For data security**, you have the option to **encrypt output files** and **decrypt input files**. Learn more about [user-controlled data encryption](/knowledgebase/user-controlled-encryption). ## PDF.co Request size API requests do not support request sizes of more than `4` megabytes in size. Please ensure that request sizes do not exceed this limit. # Webhook and Callbacks Source: https://developer.pdf.co/api/webhooks Configure callback URLs and webhooks to receive PDF.co job completion events, including timeout behavior, retry logic, and Basic Authentication. Webhooks and callbacks are both mechanisms used to enable communication between different systems or components in software development, often for the purpose of executing code in response to specific events. However, they operate in different contexts and are used in different ways. ## What are Callbacks? Callbacks are functions that are passed as arguments to other functions or methods, and they are invoked (or "called back") after the completion of a task or operation. Callbacks are a common pattern in programming, particularly in *asynchronous* operations, where you want to execute a piece of code after a certain task is done without blocking the main execution thread. ## What are Webhooks? Webhooks are a type of callback that operates over the web, allowing one system to send real-time data to another system as soon as an event occurs. Unlike traditional callbacks, which are typically defined within the same codebase or application, webhooks are used for communication between different applications or services over the internet. ## PDF.co & Webhooks The webhook or callback endpoint must respond immediately within a few seconds. If the endpoint does not respond within **20-30 seconds**, the PDF.co API considers the attempt failed and will retry the callback up to **3 times**. Webhooks can be utilized by setting the `callback` object with the URL of your choice, which triggers a `POST` request to the specified webhook URL upon completion of the job process. ```json theme={null} { "callback": "https://example.com/callback/url/you/provided" } ``` ### Basic Authentication in Callback URLs You can specify Basic Authentication credentials directly in the callback URL like this: ```json theme={null} { "callback": "https://:@example.com/test/" } ``` If your password contains special characters (like `@`, `:`, or `/`), be sure to [URL encode](https://developer.mozilla.org/en-US/docs/Glossary/percent-encoding) them. #### Example (with encoded password) ```json theme={null} { "callback": "https://myuser:my%40password@example.com/test/" } ``` This is useful when the callback endpoint requires HTTP Basic Authentication. The following APIs accept webhooks: | Controller | Endpoint | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `AI Invoice Parser Controller` | [ai-invoice-parser](/api/ai-invoice-parser) | | `Barcode Controller` | [barcode/generate](/api/barcode/generate), [barcode/read/from/url](/api/barcode/read) | | `Convert to PDF Controller` | [xls/convert/to/pdf](/api/convert-from-excel/pdf), [pdf/convert/from/csv](/api/pdf-from-document/csv), [pdf/convert/from/doc](/api/pdf-from-document/doc), [pdf/convert/from/html](/api/pdf-from-html/convert), [pdf/convert/from/image](/api/pdf-from-image), [pdf/convert/from/url](/api/pdf-from-url), [pdf/convert/from/email](/api/pdf-from-email) | | `Document Parser Controller` | [pdf/documentparser](/api/documentparser/parser) | | `Email Controller` | [email/send](/api/email/send), [email/extract-attachments](/api/email/extract-attachments), [email/decode](/api/email/decode) | | `PDF Controller` | [pdf/merge](/api/merge/pdf), [pdf/merge2](/api/merge/various-files), [pdf/split](/api/pdf-split/by-pages), [pdf/split2](/api/pdf-split/by-text-search-or-barcode), [pdf/info](/api/pdf-info-reader), [pdf/info/fields](/api/forms/info-reader), [pdf/find](/api/pdf-find/basic), [pdf/find/table](/api/pdf-find/table), [pdf/security/add](/api/pdf-password/add), [pdf/security/remove](/api/pdf-password/remove), [pdf/classifier](/api/document-classifier), [pdf/attachments/extract](/api/email/extract-attachments) | | `PDF Data Extraction Controller` | [pdf/convert/to/csv](/api/pdf-to-csv), [pdf/convert/to/html](/api/pdf-to-html), [pdf/convert/to/json2](/api/pdf-to-json/with-ai), [pdf/convert/to/text](/api/pdf-to-text/basic), [pdf/convert/to/text-simple](/api/pdf-to-text/simple), [pdf/convert/to/xls](/api/pdf-to-excel/xls), [pdf/convert/to/xlsx](/api/pdf-to-excel/xlsx), [pdf/convert/to/xml](/api/pdf-to-xml) | | `PDF Edit Controller` | [pdf/makesearchable](/api/pdf-change-text-searchable/searchable), [pdf/makeunsearchable](/api/pdf-change-text-searchable/unsearchable), [pdf/edit/add](/api/pdf-add), [pdf/edit/rotate](/api/pdf-rotate/basic), [pdf/edit/rotate/auto](/api/pdf-rotate/auto), [pdf/edit/delete-pages](/api/pdf-delete-pages), [pdf/edit/replace-text](/api/pdf-search-text-and-replace/text), [pdf/edit/delete-text](/api/pdf-search-text-and-delete), [pdf/edit/replace-text-with-image](/api/pdf-search-text-and-replace/image) | | `PDF to Image Controller` | [pdf/convert/to/jpg](/api/pdf-to-image/jpg), [pdf/convert/to/png](/api/pdf-to-image/png), [pdf/convert/to/webp](/api/pdf-to-image/webp), [pdf/convert/to/tiff](/api/pdf-to-image/tiff) | | `XLS Controller` | [xls/convert/to/csv](/api/convert-from-excel/csv), [xls/convert/to/html](/api/convert-from-excel/html), [xls/convert/to/json](/api/convert-from-excel/json), [xls/convert/to/txt](/api/convert-from-excel/text), [xls/convert/to/xml](/api/convert-from-excel/xml) | ## Summary Callbacks are functions passed into other functions to be executed after a task is completed, commonly used within the same codebase. Webhooks are a type of callback that operates over the web, allowing one system to send data to another system when a specific event occurs, typically via an HTTP POST request. In essence, **webhooks are a practical implementation of the callback concept** across different systems, facilitating real-time communication and integration between web applications. # Changelog Source: https://developer.pdf.co/changelog Notable customer-facing improvements to PDF.co, including new features, API updates, and fixes. This changelog highlights notable customer-facing improvements to PDF.co. Routine maintenance and internal infrastructure changes are not included. * Improved AI Invoice Parser job tracking, timeout handling, and result delivery. * Improved API reliability under high concurrency and repeated job-status requests. * Added clearer validation and error responses for oversized input downloads. * Changed API rate-limit windows from per-second to per-minute for smoother request handling. * Improved AI Invoice Parser result delivery and callback reliability. * Fixed PDF compression failures involving shared images. * Improved checkbox flattening and filled-rectangle rendering in PDF processing. * Improved HTML-to-PDF validation so invalid input and template errors return clearer client errors. * Improved security protections for HTML-to-PDF conversion. * Improved HTML-to-PDF support for single-page applications that use URL fragments. * Improved web-font loading reliability during HTML-to-PDF conversion. * Added OCR mode selection support to the PDF Split API. * Improved AI Invoice Parser output consistency for requested custom fields and structured line items. * Fixed duplicate material numbers caused by adjacent PDF text blocks. * Improved handling of missing and unrequested custom fields in AI Invoice Parser results. * Improved AI Invoice Parser extraction for long and multi-page invoices. * Fixed cases where structured invoice fields could be mixed into unrelated notes. * Improved timeout and job-cancellation handling for PDF conversion. * Improved error handling when converting images from remote URLs. * Added line-item structure hints to AI Invoice Parser. * Added support for JSON objects in configurable AI Invoice Parser fields. * Added page counts to AI Invoice Parser responses. * Improved validation and normalization of AI Invoice Parser line-item structures. * Improved timeout and aborted-job reporting for asynchronous processing. * Improved AI Invoice Parser extraction quality, caching, and custom-field handling. * Improved temporary-file handling and reliability for asynchronous job results. * Added clearer rate-limit details to API error responses. * Added more flexible result-retention periods and safer uploaded filename handling. * Added a legacy rendition option and configurable pre-render delay for HTML and URL conversion. * Added redaction support to PDF search-and-replace operations. * Added password support for PDF compression and information extraction. * Added Data Matrix barcode generation. * Added support for importing email attachments into PDF documents. * Improved webhook callback retries and conversion error messages. * Improved HTML-to-PDF performance and compatibility with a newer browser-rendering engine. * Improved compatibility with legacy Excel files during Excel-to-PDF conversion. * Added automatic credit refill support. * Added PDF Split and PDF Info processing to the newer processing platform. * Added TIFF and CMYK image support. * Added Google Drive downloads and presigned file-upload support. * Added profile support to simple PDF-to-text conversion. * Improved AI Invoice Parser API validation and result handling. * Added the v2 PDF compression API. * Expanded PDF editing with text, images, form fields, checkboxes, radio buttons, and formatting options. * Added profile support to HTML and URL conversion. * Added PDF flattening, crop-box handling, and PDF encryption and decryption. * Added a simple PDF-to-text API with configurable line endings. * Improved the scalability and reliability of HTML and URL conversion. * Improved credit refunds for failed and aborted jobs. * Added batch handling for pay-as-you-go billing. * Changed webhook processing so callback delivery does not consume additional credits. * Added Data Store APIs and webhook support for stored data. * Added the Delay API for scheduled processing workflows. * Expanded integration-key support across account, file, and job APIs. * Added monthly credit-pack support. * Improved optional callbacks and callbacks for aborted jobs. * Introduced the AI Invoice Parser API and Document Parser tools. * Added API management for reusable HTML templates. * Added Google Drive file support to document-processing workflows. * Added duration and job-duration fields to API responses. * Added profile support to simple PDF-to-text conversion. * Added the PDF Inspector as an alias for the PDF editing helper. * Added rotation-angle support for text watermarks and PDF editing. * Improved the Document Parser template editor and related APIs. * Improved PDF-to-image, HTML-to-PDF, and table-detection reliability. * Added TextSense document-processing integration. * Added output-link expiration information to API responses. * Improved OCR and CSV conversion behavior. * Improved PDF-to-image rendering for logos and object backgrounds. * Added a general webhook interface for processing callbacks. * Improved the dashboard, API Logs page, subscription pages, and local-time display. * Expanded PDF form-field reading and editing for checkboxes, combo boxes, and list boxes. * Added user-controlled PDF encryption and transparent colors for PDF editing. * Added direct-download link support for online file-sharing services. * Added paper-size support to image-to-PDF conversion. * Improved PDF text replacement, output filenames, and embedded-document handling. * Added public links to the PDF.co service status page and feature-request portal. * Added barcode-based PDF splitting. * Added profile support to PDF optimization. * Improved right-to-left text search and PDF text extraction. * Expanded Email Send API options and attachment information. * Added encrypted file uploads. * Added the Document Classifier and improved the Document Parser editor. * Expanded HTML-to-PDF profile controls for input styles. * Improved PDF information extraction for protected and damaged documents. # Welcome to PDF.co Docs Source: https://developer.pdf.co/index Your all-in-one guide to automate PDF, documents, and data extraction with APIs and integrations. ## Quickstarts Whether you’re a developer or a no-code builder, you’ll find examples, quickstarts and step-by-step guides to get you up and running fast. ### For API Developers ### For Automation Users ### Basic Knowledge ## Popular Endpoints Extract structured data from any invoice automatically—no templates needed. Add text, images, forms, other PDFs, fill forms, links to external sites and external PDF files. Optimize PDF files up to 13 times smaller in file size by optimizing images and objects. Extract data from PDFs, JPGs, and PNGs — including fields, tables, values, and barcodes. # Airtable Source: https://developer.pdf.co/integrations/airtable [Airtable](https://airtable.com/) is a cloud\-based project management system that was founded on the belief that software shouldn’t dictate how you work. ## Airtable and PDF.co Plugins ## Airtable and PDF.co Integration via Zapier ## How to Generate PDF Files