PDF.co ACTION
PDF Split Text Search
Split PDF by text search. See docs here
- Action
- Read only
- API key
- SDK
- MCP
IMPLEMENTATION
Call this tool
Connect a user's PDF.co account once, then configure and run PDF Split Text Search from your backend or agent.
import { PipedreamClient } from "@pipedream/sdk"
const pd = new PipedreamClient({
projectId: process.env.PIPEDREAM_PROJECT_ID!,
clientId: process.env.PIPEDREAM_CLIENT_ID!,
clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
projectEnvironment: "production",
})
const result = await pd.actions.run({
id: "pdf_co-pdf-split-by-text-search",
externalUserId: "{external_user_id}", // any stable ID for this user in your system
configuredProps: {
pdf_co: { authProvisionId: "apn_xxxxxxx" },
url: "URL",
searchString: "Search String",
},
})
console.log(result)curl -X POST https://api.pipedream.com/v1/connect/{project_id}/actions/run \
-H "Content-Type: application/json" \
-H "X-PD-Environment: production" \
-H "Authorization: Bearer {access_token}" \
-d '{
"external_user_id": "{external_user_id}",
"id": "pdf_co-pdf-split-by-text-search",
"configured_props": {
"pdf_co": { "authProvisionId": "apn_xxxxxxx" },
"url": "URL",
"searchString": "Search String"
}
}'// accessToken: mint a short-lived token with the Connect SDK — see the MCP guide
const transport = new StreamableHTTPClientTransport(
new URL("https://remote.mcp.pipedream.net/v3"),
{
requestInit: {
headers: {
Authorization: `Bearer ${accessToken}`,
"x-pd-project-id": "{project_id}",
"x-pd-environment": "production",
"x-pd-external-user-id": "{external_user_id}", // any stable ID for this user in your system
"x-pd-app-slug": "pdf_co",
},
},
},
)
const mcp = new Client({ name: "my-agent", version: "1.0.0" })
await mcp.connect(transport)
const { tools } = await mcp.listTools()
// listTools() hands your model this tool's input schema, so it can
// fill the arguments itself:
const result = await mcp.callTool({
name: "pdf_co-pdf-split-by-text-search",
arguments: {
url: "URL",
searchString: "Search String",
},
})SCHEMA
Inputs
Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.
| Property | Type | Description |
|---|---|---|
url URL | string | URL to the source file. Supports links from Google Drive, Dropbox and from built-in PDF.co files storage. Required |
searchString Search String | string | Text to search for on pages. Required |
excludeKeyPages Exclude Key Pages | boolean | Set to true if you want to exclude pages where text was found. false by default. Optional |
regexSearch Regex Search | boolean | Set to true to enable regular expressions for search string. false by default. Optional |
caseSensitive Case Sensitive | boolean | Defines if keywords in rules are case sensitive or not. Optional |
lang Lang | string | Sets language for OCR (text from image) to use for scanned PDF, PNG, JPG documents input when extracting text. Default is “eng”. Other languages are also supported: deu, spa, chi_sim, jpn and many others (full list of supported OCR languages is (here)[https://apidocs.pdf.co/kb/OCR/list-of-supported-languages-for-ocr]. You can also use 2 languages simultaneously like this: eng+deu or jpn+kor (any combination). Optional |
httpusername HTTP Username | string | HTTP auth user name if required to access source url. Optional |
httppassword HTTP Password | string | HTTP auth password if required to access source url. Optional |
name Name | string | File name for generated output. Optional |
expiration Expiration | integer | Output link expiration in minutes. Default is 60 (i.e. 60 minutes or 1 hour). Optional |
inline Inline | boolean | Must be one of: true to return data as inline or false to return link to output file (default). Optional |
async Async | boolean | Runs processing asynchronously. Returns JobId. Optional |
profiles Profiles | any | Use this parameter to set additional configuration for fine tuning and extra options. Explore PDF.co for profile examples. Optional |
REFERENCE
Tool details
Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.
- Registry key
- pdf_co-pdf-split-by-text-search
- Version
- 0.0.2
- App
- PDF.co
- Authentication
- API key
- Read-only
- Yes
- Destructive
- No
- Open world
- Yes
- Source
- View on GitHub ↗