PDF.co ACTION
Document Parser
Document Parser can automatically parse PDF, JPG, PNG document to extract fields, tables, values, barcodes from invoices, statements, orders and other PDF and scanned documents. See docs here
- Action
- Read only
- API key
- SDK
- MCP
IMPLEMENTATION
Call this tool
Connect a user's PDF.co account once, then configure and run Document Parser from your backend or agent.
import { PipedreamClient } from "@pipedream/sdk"
const pd = new PipedreamClient({
projectId: process.env.PIPEDREAM_PROJECT_ID!,
clientId: process.env.PIPEDREAM_CLIENT_ID!,
clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
projectEnvironment: "production",
})
const result = await pd.actions.run({
id: "pdf_co-document-parser",
externalUserId: "{external_user_id}", // any stable ID for this user in your system
configuredProps: {
pdf_co: { authProvisionId: "apn_xxxxxxx" },
url: "URL",
outputFormat: "Output Format",
},
})
console.log(result)curl -X POST https://api.pipedream.com/v1/connect/{project_id}/actions/run \
-H "Content-Type: application/json" \
-H "X-PD-Environment: production" \
-H "Authorization: Bearer {access_token}" \
-d '{
"external_user_id": "{external_user_id}",
"id": "pdf_co-document-parser",
"configured_props": {
"pdf_co": { "authProvisionId": "apn_xxxxxxx" },
"url": "URL",
"outputFormat": "Output Format"
}
}'// accessToken: mint a short-lived token with the Connect SDK — see the MCP guide
const transport = new StreamableHTTPClientTransport(
new URL("https://remote.mcp.pipedream.net/v3"),
{
requestInit: {
headers: {
Authorization: `Bearer ${accessToken}`,
"x-pd-project-id": "{project_id}",
"x-pd-environment": "production",
"x-pd-external-user-id": "{external_user_id}", // any stable ID for this user in your system
"x-pd-app-slug": "pdf_co",
},
},
},
)
const mcp = new Client({ name: "my-agent", version: "1.0.0" })
await mcp.connect(transport)
const { tools } = await mcp.listTools()
// listTools() hands your model this tool's input schema, so it can
// fill the arguments itself:
const result = await mcp.callTool({
name: "pdf_co-document-parser",
arguments: {
url: "URL",
outputFormat: "Output Format",
},
})SCHEMA
Inputs
Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.
| Property | Type | Description |
|---|---|---|
url URL | string | URL to the source file. Supports links from Google Drive, Dropbox and from built-in PDF.co files storage. Required |
outputFormat Output Format | string | Default is JSON. You can override default output format to CSV or XML to generate CSV or XML output accordingly. Optional |
httpusername HTTP Username | string | HTTP auth user name if required to access source url. Optional |
httppassword HTTP Password | string | HTTP auth password if required to access source url. Optional |
password PDF Password | string | Password of PDF file. Optional |
inline Inline | boolean | Must be one of: true to return data as inline or false to return link to output file (default). Optional |
name Name | string | File name for generated output. Optional |
expiration Expiration | integer | Output link expiration in minutes. Default is 60 (i.e. 60 minutes or 1 hour). Optional |
async Async | boolean | Runs processing asynchronously. Returns JobId. Optional |
profiles Profiles | any | Use this parameter to set additional configuration for fine tuning and extra options. Explore PDF.co for profile examples. Optional |
REFERENCE
Tool details
Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.
- Registry key
- pdf_co-document-parser
- Version
- 0.0.3
- App
- PDF.co
- Authentication
- API key
- Read-only
- Yes
- Destructive
- No
- Open world
- Yes
- Source
- View on GitHub ↗