View as Markdown
Databricks icon

Databricks ACTION

Create Job

Create a job. See the documentation
  • Action
  • Writes data
  • API key
  • SDK
  • MCP

IMPLEMENTATION

Call this tool

Connect a user's Databricks account once, then configure and run Create Job from your backend or agent.

import { PipedreamClient } from "@pipedream/sdk"

const pd = new PipedreamClient({
  projectId: process.env.PIPEDREAM_PROJECT_ID!,
  clientId: process.env.PIPEDREAM_CLIENT_ID!,
  clientSecret: process.env.PIPEDREAM_CLIENT_SECRET!,
  projectEnvironment: "production",
})

const result = await pd.actions.run({
  id: "databricks-create-job",
  externalUserId: "{external_user_id}", // any stable ID for this user in your system
  configuredProps: {
    databricks: { authProvisionId: "apn_xxxxxxx" },
    tasks: ["Tasks"],
    name: "Job Name",
  },
})

console.log(result)

SCHEMA

Inputs

Pipedream supplies the connected account. Your application provides the operation-specific values below. Dynamic inputs are resolved against that user's account.

Create Job inputs
Property Type Description
tasks Tasks string[]

A list of task specifications to be executed by this job. JSON string format. See the API documentation for task specification details.

Example:

[
  {
    "notebook_task": {
      "notebook_path": "/Workspace/Users/sharky@databricks.com/weather_ingest"
    },
    "task_key": "weather_ocean_data"
  }
]
Required
name Job Name string
An optional name for the job
Optional
tags Tags object
A map of tags associated with the job. These are forwarded to the cluster as cluster tags for jobs clusters, and are subject to the same limitations as cluster tags
Optional
jobClusters Job Clusters string[]

A list of job cluster specifications that can be shared and reused by tasks of this job. JSON string format. See the API documentation for job cluster specification details.

Example:

[
  {
    "job_cluster_key": "auto_scaling_cluster",
    "new_cluster": {
      "autoscale": {
        "max_workers": 16,
        "min_workers": 2
      },
      "node_type_id": null,
      "spark_conf": {
        "spark.speculation": true
      },
      "spark_version": "7.3.x-scala2.12"
    }
  }
]
Optional
emailNotifications Email Notifications string

An optional set of email addresses to notify when runs of this job begin, complete, or when the job is deleted. Specify as a JSON object with keys for each notification type. See the API documentation for details on each field.

Example:

{
  "on_start": ["user1@example.com"],
  "on_success": ["user2@example.com"],
  "on_failure": ["user3@example.com"],
  "on_duration_warning_threshold_exceeded": ["user4@example.com"],
  "on_streaming_backlog_exceeded": ["user5@example.com"]
}
Optional
webhookNotifications Webhook Notifications string

A collection of system notification IDs to notify when runs of this job begin, complete, or encounter specific events. Specify as a JSON object with keys for each notification type. Each key accepts an array of objects with an id property (system notification ID). A maximum of 3 destinations can be specified for each property.

Supported keys:

  • on_start: Notified when the run starts.
  • on_success: Notified when the run completes successfully.
  • on_failure: Notified when the run fails.
  • on_duration_warning_threshold_exceeded: Notified when the run duration exceeds the specified threshold.
  • on_streaming_backlog_exceeded: Notified when streaming backlog thresholds are exceeded.

See the API documentation for details.

Example:

{
  "on_success": [
    { "id": "https://eoiqkb8yzox6u2n.m.pipedream.net" }
  ],
  "on_failure": [
    { "id": "https://another-webhook-url.com/notify" }
  ]
}
Optional
timeoutSeconds Timeout Seconds integer
An optional timeout applied to each run of this job. The default behavior is to have no timeout
Optional
schedule Schedule string

An optional periodic schedule for this job, specified as a JSON object. By default, the job only runs when triggered manually or via the API. The schedule object must include:

  • quartz_cron_expression (required): A Cron expression using Quartz syntax that defines when the job runs. See Cron Trigger details.
  • timezone_id (required): A Java timezone ID (e.g., "Europe/London") that determines the timezone for the schedule. See Java TimeZone details.
  • pause_status (optional): Set to "UNPAUSED" (default) or "PAUSED" to control whether the schedule is active.

Example:

{
  "quartz_cron_expression": "0 0 12 * * ?",
  "timezone_id": "Asia/Ho_Chi_Minh",
  "pause_status": "UNPAUSED"
}
Optional
maxConcurrentRuns Max Concurrent Runs integer
An optional maximum allowed number of concurrent runs of the job. Defaults to 1
Optional
gitSource Git Source string

An optional specification for a remote Git repository containing the source code used by tasks. Provide as a JSON string.

This enables version-controlled source code for notebook, dbt, Python script, and SQL File tasks. If git_source is set, these tasks retrieve files from the remote repository by default (can be overridden per task by setting source to WORKSPACE). Note: dbt and SQL File tasks require git_source to be defined. See the API documentation for more details.

Fields:

  • git_url (required): URL of the repository to be cloned (e.g., "https://github.com/databricks/databricks-cli").
  • git_provider (required): Service hosting the repository. One of: gitHub, bitbucketCloud, azureDevOpsServices, gitHubEnterprise, bitbucketServer, gitLab, gitLabEnterpriseEdition, awsCodeCommit.
  • git_branch: Name of the branch to check out (cannot be used with git_tag or git_commit).
  • git_tag: Name of the tag to check out (cannot be used with git_branch or git_commit).
  • git_commit: Commit hash to check out (cannot be used with git_branch or git_tag).

Example:

{
  "git_url": "https://github.com/databricks/databricks-cli",
  "git_provider": "gitHub",
  "git_branch": "main"
}
Optional
accessControlList Access Control List string[]

A list of permissions to set on the job, specified as a JSON array of objects. Each object can define permissions for a user, group, or service principal.

Each object may include:

  • user_name: Name of the user.
  • group_name: Name of the group.
  • service_principal_name: Application ID of a service principal.
  • permission_level: Permission level. One of: CAN_MANAGE, IS_OWNER, CAN_MANAGE_RUN, CAN_VIEW.

Example:

[
  {
    "permission_level": "IS_OWNER",
    "user_name": "jorge.c@turing.com"
  },
  {
    "permission_level": "CAN_VIEW",
    "group_name": "data-scientists"
  }
]

See the API documentation for more details.

Optional

REFERENCE

Tool details

Behavior hints are published with the component in the Pipedream registry and surface as MCP tool annotations, so an agent can reason about a tool before it calls it.

Registry key
databricks-create-job
Version
0.0.4
App
Databricks
Authentication
API key
Read-only
No
Destructive
No
Open world
Yes