View as Markdown
Scrapfly icon

Scrapfly ACTION

Scrape Page

Extract data from a specified web page. See the documentation
  • Action
  • Read only
  • API key
  • SDK
  • MCP

IMPLEMENTATION

Call this tool

Connect a user's Scrapfly account once, then configure and run Scrape Page from your backend or agent.

import { PipedreamClient } from "@pipedream/sdk"

const pd = new PipedreamClient({
  projectId: process.env.PIPEDREAM_PROJECT_ID!,
  clientId: process.env.PIPEDREAM_CLIENT_ID!,
  clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
  projectEnvironment: "production",
})

const result = await pd.actions.run({
  id: "scrapfly-scrape-page",
  externalUserId: "{external_user_id}", // any stable ID for this user in your system
  configuredProps: {
    scrapfly: { authProvisionId: "apn_xxxxxxx" },
    url: "URL",
    headers: "Headers",
  },
})

console.log(result)

SCHEMA

Inputs

Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.

Scrape Page inputs
Property Type Description
url URL string
This URL is used to transform any relative URLs in the document into absolute URLs automatically. It can be either the base URL or the exact URL of the document. Must be url encoded.
Required
headers Headers object
Pass custom headers to the request.
Optional
lang Language string
Select page language. By default it uses the language of the selected proxy location. Behind the scenes, it configures the Accept-Language HTTP header. If the website support the language, the content will be in that lang. Note: you cannot set headers Accept-Language header manually. See the documentation
Optional
os Operating System string
Operating System, if not selected it's random. Note: you cannot set os parameter and User-Agent header at the same time. See the documentation
Optional
timeout Timeout integer
Timeout in milliseconds. It represents the maximum time allowed for Scrapfly to perform the scrape. Since timeout is not trivial to understand see our extended documentation on timeouts
Optional
format Format string
Format of the response.
Optional
retry Retry boolean
Improve reliability with retries on failure.
Optional
proxifiedResponse Proxified Response boolean
Return the content of the page directly.
Optional
debug Debug boolean
Store the API result and take a screenshot if rendering js is enabled.
Optional
correlationId Correlation ID string
Helper ID for correlating a group of scrapes.
Optional
tags Tags string[]
Add tags to your scrapes to group them.
Optional
dns DNS boolean
Query and retrieve target DNS information.
Optional
ssl SSL boolean
SSL option.
Optional
proxyPool Proxy Pool string
Select the proxy pool to use.
Optional

REFERENCE

Tool details

Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.

Registry key
scrapfly-scrape-page
Version
0.0.2
App
Scrapfly
Authentication
API key
Read-only
Yes
Destructive
No
Open world
Yes