# Scrape Website — Webscrape AI

> Scrape the provided URL and store the results in the system. See the documentation

- Key: `webscrape_ai-scrape-website`
- Type: Action (Write)
- Version: 0.0.3
- App: Webscrape AI (`webscrape_ai`) — https://pipedream.com/apps/webscrape-ai.md
- This page (HTML): https://pipedream.com/apps/webscrape-ai/actions/scrape-website
- Hints: open-world
- Source: https://github.com/PipedreamHQ/pipedream/blob/master/components/webscrape_ai/actions/scrape-website/scrape-website.mjs

## Description

Scrape the provided URL and store the results in the system. [See the documentation](https://webscrapeai.com/docs)

## Props

| Prop | Type | Required | Description |
|---|---|---|---|
| `url` | `string` | Yes | The URL of the website to scrape |
| `command` | `string` | Yes | The data you want to extract. E.g. I want to extract all the news details |
| `schema` | `string` | Yes | Schema representing the fields you want to scrape. E.g. {"author":"string","comments_count":"integer","points":"integer","posted_time":"string","title":"string","url":"url"} |
| `pages` | `integer` | No | Number of pages to scrape. Default value is 1. |
| `headers` | `string` | No | List of headers in key-value pairs. i.e Accept: application/json |
| `instructions` | `string` | No | List of JavaScript instructions that you want to execute, like clicking a specific button, waiting for a specific code block to appear, etc. Example: {"click": "#button_id"}. See the documentation for more information. |

## Run it

**MCP**

```ts
import { Client } from "@modelcontextprotocol/sdk/client/index.js"
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js"
import { PipedreamClient } from "@pipedream/sdk"

const pd = new PipedreamClient({
  projectId: process.env.PIPEDREAM_PROJECT_ID!,
  clientId: process.env.PIPEDREAM_CLIENT_ID!,
  clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
  projectEnvironment: "production",
})

const accessToken = await pd.rawAccessToken

const transport = new StreamableHTTPClientTransport(
  new URL("https://remote.mcp.pipedream.net/v3"),
  {
    requestInit: {
      headers: {
        Authorization: `Bearer ${accessToken}`,
        "x-pd-project-id": process.env.PIPEDREAM_PROJECT_ID!,
        "x-pd-environment": "production",
        "x-pd-external-user-id": "{external_user_id}", // any stable ID for this user in your system
        "x-pd-app-slug": "webscrape_ai",
      },
    },
  },
)

const mcp = new Client({ name: "my-agent", version: "1.0.0" })
await mcp.connect(transport)

const { tools } = await mcp.listTools()

// listTools() hands your model this tool's input schema, so it can
// fill the arguments itself:
const result = await mcp.callTool({
  name: "webscrape_ai-scrape-website",
  arguments: {
    url: "URL",
    command: "Command",
  },
})
```

**TypeScript**

```ts
import { PipedreamClient } from "@pipedream/sdk"

const pd = new PipedreamClient({
  projectId: process.env.PIPEDREAM_PROJECT_ID!,
  clientId: process.env.PIPEDREAM_CLIENT_ID!,
  clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
  projectEnvironment: "production",
})

const result = await pd.actions.run({
  id: "webscrape_ai-scrape-website",
  externalUserId: "{external_user_id}", // any stable ID for this user in your system
  configuredProps: {
    webscrape_ai: { authProvisionId: "apn_xxxxxxx" },
    url: "URL",
    command: "Command",
  },
})

console.log(result)
```

**cURL**

```bash
curl -X POST https://api.pipedream.com/v1/connect/{project_id}/actions/run \
  -H "Content-Type: application/json" \
  -H "X-PD-Environment: production" \
  -H "Authorization: Bearer {access_token}" \
  -d '{
    "external_user_id": "{external_user_id}",
    "id": "webscrape_ai-scrape-website",
    "configured_props": {
      "webscrape_ai": { "authProvisionId": "apn_xxxxxxx" },
      "url": "URL",
      "command": "Command"
    }
  }'
```

---

- App: https://pipedream.com/apps/webscrape-ai.md · All apps: https://pipedream.com/apps
