View as Markdown
PDF.co icon

PDF.co ACTION

PDF Split Text Search

Split PDF by text search. See docs here
  • Action
  • Read only
  • API key
  • SDK
  • MCP

IMPLEMENTATION

Call this tool

Connect a user's PDF.co account once, then configure and run PDF Split Text Search from your backend or agent.

import { PipedreamClient } from "@pipedream/sdk"

const pd = new PipedreamClient({
  projectId: process.env.PIPEDREAM_PROJECT_ID!,
  clientId: process.env.PIPEDREAM_CLIENT_ID!,
  clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
  projectEnvironment: "production",
})

const result = await pd.actions.run({
  id: "pdf_co-pdf-split-by-text-search",
  externalUserId: "{external_user_id}", // any stable ID for this user in your system
  configuredProps: {
    pdf_co: { authProvisionId: "apn_xxxxxxx" },
    url: "URL",
    searchString: "Search String",
  },
})

console.log(result)

SCHEMA

Inputs

Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.

PDF Split Text Search inputs
Property Type Description
url URL string
URL to the source file. Supports links from Google Drive, Dropbox and from built-in PDF.co files storage.
Required
searchString Search String string
Text to search for on pages.
Required
excludeKeyPages Exclude Key Pages boolean
Set to true if you want to exclude pages where text was found. false by default.
Optional
regexSearch Regex Search boolean
Set to true to enable regular expressions for search string. false by default.
Optional
caseSensitive Case Sensitive boolean
Defines if keywords in rules are case sensitive or not.
Optional
lang Lang string
Sets language for OCR (text from image) to use for scanned PDF, PNG, JPG documents input when extracting text. Default is “eng”. Other languages are also supported: deu, spa, chi_sim, jpn and many others (full list of supported OCR languages is (here)[https://apidocs.pdf.co/kb/OCR/list-of-supported-languages-for-ocr]. You can also use 2 languages simultaneously like this: eng+deu or jpn+kor (any combination).
Optional
httpusername HTTP Username string
HTTP auth user name if required to access source url.
Optional
httppassword HTTP Password string
HTTP auth password if required to access source url.
Optional
name Name string
File name for generated output.
Optional
expiration Expiration integer
Output link expiration in minutes. Default is 60 (i.e. 60 minutes or 1 hour).
Optional
inline Inline boolean
Must be one of: true to return data as inline or false to return link to output file (default).
Optional
async Async boolean
Runs processing asynchronously. Returns JobId.
Optional
profiles Profiles any
Use this parameter to set additional configuration for fine tuning and extra options. Explore PDF.co for profile examples.
Optional

REFERENCE

Tool details

Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.

Registry key
pdf_co-pdf-split-by-text-search
Version
0.0.2
App
PDF.co
Authentication
API key
Read-only
Yes
Destructive
No
Open world
Yes