Dataiku ACTION
Build Dataset
Start a job that builds one or more outputs (typically datasets) in a DSS project. Use this to rebuild specific outputs directly; use Run Scenario instead when the pipeline is already orchestrated as a scenario. Use List Datasets to find valid output names. A successful call only means the job was accepted — the response's
id is the job ID, which you pass to Get Job Status to follow it to completion. Requires the RUN_JOBS privilege on the project. See the documentation- Action
- Writes data
- API key
- SDK
- MCP
IMPLEMENTATION
Call this tool
Connect a user's Dataiku account once, then configure and run Build Dataset from your backend or agent.
import { PipedreamClient } from "@pipedream/sdk"
const pd = new PipedreamClient({
projectId: process.env.PIPEDREAM_PROJECT_ID!,
clientId: process.env.PIPEDREAM_CLIENT_ID!,
clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
projectEnvironment: "production",
})
const result = await pd.actions.run({
id: "dataiku-build-dataset",
externalUserId: "{external_user_id}", // any stable ID for this user in your system
configuredProps: {
dataiku: { authProvisionId: "apn_xxxxxxx" },
projectKey: "Project Key",
outputIds: ["Outputs To Build"],
},
})
console.log(result)curl -X POST https://api.pipedream.com/v1/connect/{project_id}/actions/run \
-H "Content-Type: application/json" \
-H "X-PD-Environment: production" \
-H "Authorization: Bearer {access_token}" \
-d '{
"external_user_id": "{external_user_id}",
"id": "dataiku-build-dataset",
"configured_props": {
"dataiku": { "authProvisionId": "apn_xxxxxxx" },
"projectKey": "Project Key",
"outputIds": ["Outputs To Build"]
}
}'// accessToken: mint a short-lived token with the Connect SDK — see the MCP guide
const transport = new StreamableHTTPClientTransport(
new URL("https://remote.mcp.pipedream.net/v3"),
{
requestInit: {
headers: {
Authorization: `Bearer ${accessToken}`,
"x-pd-project-id": "{project_id}",
"x-pd-environment": "production",
"x-pd-external-user-id": "{external_user_id}", // any stable ID for this user in your system
"x-pd-app-slug": "dataiku",
},
},
},
)
const mcp = new Client({ name: "my-agent", version: "1.0.0" })
await mcp.connect(transport)
const { tools } = await mcp.listTools()
// listTools() hands your model this tool's input schema, so it can
// fill the arguments itself:
const result = await mcp.callTool({
name: "dataiku-build-dataset",
arguments: {
projectKey: "Project Key",
outputIds: ["Outputs To Build"],
},
})SCHEMA
Inputs
Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.
| Property | Type | Description |
|---|---|---|
projectKey Project Key | string | The key of the DSS project, e.g. MYPROJECT. Call List Projects and pass the projectKey field of the project you want. This is an identifier, not the project's display name — a display name will not resolve. In the Dataiku DSS GUI the same value appears in the project's URL as /projects/MYPROJECT/. Required |
outputIds Outputs To Build | string[] | Names of the outputs to build, e.g. ["customers_prepared"]. Call List Datasets and pass the name field of each dataset you want built. Required |
buildType Build Type | string | How far upstream the build should go. RECURSIVE_BUILD also builds upstream dependencies that are out of date, RECURSIVE_MISSING_ONLY_BUILD only builds upstream items that do not exist yet, NON_RECURSIVE_FORCED_BUILD rebuilds just the requested outputs, and RECURSIVE_FORCED_BUILD rebuilds the requested outputs and everything upstream regardless of whether they are up to date. Optional |
partition Partition | string | The partition to build for each requested output, e.g. 2015-07-07. Omit for non-partitioned datasets. Optional |
REFERENCE
Tool details
Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.
- Registry key
- dataiku-build-dataset
- Version
- 0.0.2
- App
- Dataiku
- Authentication
- API key
- Read-only
- No
- Destructive
- No
- Open world
- Yes
- Source
- View on GitHub ↗