Extract Text with LLMWhisperer API on New Blog Posts from HubSpot API

Pipedream makes it easy to connect APIs for LLMWhisperer, HubSpot and 2,200+ other apps remarkably fast.

Trigger workflow on

New Blog Posts from the HubSpot API

Next, do this

Extract Text with the LLMWhisperer API

No credit card required

▶

Watch us build a workflow

8 min

Watch now ➜

Trusted by 1,000,000+ developers from startups to Fortune 500 companies

Developers ♥ Pipedream

Getting Started#

This integration creates a workflow with a HubSpot trigger and LLMWhisperer action. When you configure and deploy the workflow, it will run on Pipedream's servers 24x7 for free.

Select this integration
Configure the New Blog Posts trigger
1. Connect your HubSpot account
2. Configure timer
Configure the Extract Text action
1. Connect your LLMWhisperer account
2. Select a Processing Mode
3. Select a Output Mode
4. Optional- Configure Page Seperator
5. Optional- Configure Force Text Processing
6. Optional- Configure Pages To Extract
7. Optional- Configure Timeout
8. Optional- Configure Store Metadata for Highlighting
9. Optional- Configure Median Filter Size
10. Optional- Configure Gaussian Blur Radius
11. Optional- Select a OCR Provider
12. Optional- Configure Line Splitter Tolerance
13. Optional- Configure Horizontal Stretch Factor
14. Configure URL In Post
Deploy the workflow
Send a test event to validate your setup
Turn on the trigger

Details#

This integration uses pre-built, source-available components from Pipedream's GitHub repo. These components are developed by Pipedream and the community, and verified and maintained by Pipedream.

To contribute an update to an existing component or create a new component, create a PR on GitHub. If you're new to Pipedream component development, you can start with quickstarts for trigger span and action development, and then review the component API reference.

Trigger#

New Blog Posts on HubSpot

Description:Emit new event for each new blog post.

Version:0.0.16

Key:hubspot-new-blog-article

View on GitHub

HubSpot Overview#

The HubSpot API enables developers to integrate into HubSpots CRM, CMS, Conversations, and other features. It allows for automated management of contacts, companies, deals, and marketing campaigns, enabling custom workflows, data synchronization, and task automation. This streamlines operations and boosts customer engagement, with real-time updates for rapid response to market changes.

Trigger Code#

import common from "../common/common.mjs";

export default {
  ...common,
  key: "hubspot-new-blog-article",
  name: "New Blog Posts",
  description: "Emit new event for each new blog post.",
  version: "0.0.16",
  dedupe: "unique",
  type: "source",
  hooks: {},
  methods: {
    ...common.methods,
    getTs(blogpost) {
      return Date.parse(blogpost.created);
    },
    generateMeta(blogpost) {
      const {
        id,
        name: summary,
      } = blogpost;
      const ts = this.getTs(blogpost);
      return {
        id,
        summary,
        ts,
      };
    },
    getParams(after) {
      return {
        limit: 100,
        createdAfter: after, // return entries created since event last ran
      };
    },
    async processResults(after, params) {
      await this.paginate(
        params,
        this.hubspot.getBlogPosts.bind(this),
        "results",
      );
    },
  },
};

Trigger Configuration#

This component may be configured based on the props defined in the component code. Pipedream automatically prompts for input values in the UI and CLI.

Label	Prop	Type	Description
HubSpot	`hubspot`	`app`	This component uses the HubSpot app.
N/A	`db`	`$.service.db`	This component uses `$.service.db` to maintain state between executions.
	`timer`	`$.interface.timer`

Trigger Authentication#

HubSpot uses OAuth authentication. When you connect your HubSpot account, Pipedream will open a popup window where you can sign into HubSpot and grant Pipedream permission to connect to your account. Pipedream securely stores and automatically refreshes the OAuth tokens so you can easily authenticate any HubSpot API.

Pipedream requests the following authorization scopes when you connect your account:

business-intelligencecrm.lists.readcrm.lists.writecrm.objects.companies.readcrm.objects.companies.writecrm.objects.contacts.readcrm.objects.contacts.writecrm.objects.deals.readcrm.objects.deals.writecrm.objects.quotes.readcrm.objects.quotes.writecrm.objects.owners.readcrm.schemas.companies.readcrm.schemas.companies.writecrm.schemas.contacts.readcrm.schemas.contacts.writecrm.schemas.deals.readcrm.schemas.deals.writecrm.schemas.quotes.readconversations.readcrm.importfilesformsforms-uploaded-filesintegration-syncoauthtimeline

About HubSpot#

HubSpot's CRM platform contains the marketing, sales, service, operations, and website-building software you need to grow your business.

Action#

Extract Text on LLMWhisperer

Description:Convert your PDF/scanned documents to text format which can be used by LLMs. [See the documentation](https://docs.unstract.com/llm_whisperer/apis/llm_whisperer_text_extraction_api)

Version:0.0.1

Key:llmwhisperer-extract-text

View on GitHub

Action Code#

import fs from "fs";
import app from "../../llmwhisperer.app.mjs";

export default {
  key: "llmwhisperer-extract-text",
  name: "Extract Text",
  description: "Convert your PDF/scanned documents to text format which can be used by LLMs. [See the documentation](https://docs.unstract.com/llm_whisperer/apis/llm_whisperer_text_extraction_api)",
  version: "0.0.1",
  type: "action",
  props: {
    app,
    processingMode: {
      type: "string",
      label: "Processing Mode",
      description: "The processing mode to be used. Choose between `ocr` and `text`.",
      options: [
        "ocr",
        "text",
      ],
    },
    outputMode: {
      type: "string",
      label: "Output Mode",
      description: "The output mode to be used. Choose between `line-printer` and `text`.",
      options: [
        "line-printer",
        "text",
      ],
    },
    pageSeperator: {
      type: "string",
      label: "Page Seperator",
      description: "The string to be used as a page separator. Eg: `<<<`",
      optional: true,
    },
    forceTextProcessing: {
      type: "boolean",
      label: "Force Text Processing",
      description: "If set to true, the document will be processed as text only. If set to false, the document will be processed based on LLMWhisperer's chosed stratergy.",
      optional: true,
    },
    pagesToExtract: {
      type: "string",
      label: "Pages To Extract",
      description: "Define which pages to extract. By default all pages are extracted. You can specify which pages to extract with this parameter. Example `1-5,7,21-` will extract pages **1,2,3,4,5,7,21,22,23,24...** till the last page.",
      optional: true,
    },
    timeout: {
      type: "integer",
      label: "Timeout",
      description: "The time in seconds after which the request will automatically switch to async mode. If a timeout occurs then the API will return a 202 message along with `whisper-hash` which can be used later to check processing status and retrieve the text. Refer to the async operation documentation for more information",
      optional: true,
    },
    storeMetadataForHighlighting: {
      type: "boolean",
      label: "Store Metadata for Highlighting",
      description: "If set to true, metadata required for the highlighting is stored. If you do not require highlighting API, set this to false. Note that setting this to true will store your text in our servers",
      optional: true,
    },
    medianFilterSize: {
      type: "integer",
      label: "Median Filter Size",
      description: "The size of the median filter to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
    },
    gaussianBlurRadius: {
      type: "integer",
      label: "Gaussian Blur Radius",
      description: "The radius of the gaussian blur to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
    },
    ocrProvider: {
      type: "string",
      label: "OCR Provider",
      description: "The OCR provider to be used. Choose between `simple` and `advanced`. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
      options: [
        "simple",
        "advanced",
      ],
    },
    lineSplitterTolerance: {
      type: "string",
      label: "Line Splitter Tolerance",
      description: "Factor to decide when to move text to the next line when it is above or below the baseline. The default value of `0.4` signifies 40% of the average character height.",
      optional: true,
    },
    horizontalStretchFactor: {
      type: "string",
      label: "Horizontal Stretch Factor",
      description: "Factor by which a horizontal stretch has to applied. It defaults to `1.0`. A stretch factor of `1.1` would mean at 10% stretch factor applied. Normally this factor need not be adjusted. You might want to use this parameter when multi column layouts back into each other. For example in a two column layout, the two columns get merged into one.",
      optional: true,
    },
    urlInPost: {
      type: "boolean",
      label: "URL In Post",
      description: "If set to `true`, the headers will be set to `text/plain`. If set to `false`, the headers will be set to `application/octet-stream`.",
      reloadProps: true,
      default: true,
    },
  },
  additionalProps() {
    const { urlInPost } = this;
    return {
      data: {
        type: "string",
        label: urlInPost
          ? "Document URL"
          : "Document Path",
        description: urlInPost
          ? "The URL of the document to process."
          : "Document path of the file previously downloaded in Pipedream E.g. (`/tmp/my-file.txt`). [Download a file to the `/tmp` directory](https://pipedream.com/docs/code/nodejs/http-requests/#download-a-file-to-the-tmp-directory)",
      },
    };
  },
  methods: {
    getHeaders(urlInPost) {
      return {
        "Content-Type": urlInPost
          ? "text/plain"
          : "application/octet-stream",
      };
    },
    getData(urlInPost, data) {
      return urlInPost
        ? data
        : fs.readFileSync(data);
    },
    extractText(args = {}) {
      return this.app.post({
        path: "/whisper",
        ...args,
      });
    },
  },
  async run({ $ }) {
    const {
      extractText,
      getHeaders,
      getData,
      urlInPost,
      processingMode,
      outputMode,
      pageSeperator,
      forceTextProcessing,
      pagesToExtract,
      timeout,
      storeMetadataForHighlighting,
      medianFilterSize,
      gaussianBlurRadius,
      ocrProvider,
      lineSplitterTolerance,
      horizontalStretchFactor,
      data,
    } = this;

    const response = await extractText({
      $,
      headers: getHeaders(urlInPost),
      params: {
        url_in_post: urlInPost,
        processing_mode: processingMode,
        output_mode: outputMode,
        page_seperator: pageSeperator,
        force_text_processing: forceTextProcessing,
        pages_to_extract: pagesToExtract,
        timeout,
        store_metadata_for_highlighting: storeMetadataForHighlighting,
        median_filter_size: medianFilterSize,
        gaussian_blur_radius: gaussianBlurRadius,
        ocr_provider: ocrProvider,
        line_splitter_tolerance: lineSplitterTolerance,
        horizontal_stretch_factor: horizontalStretchFactor,
      },
      data: getData(urlInPost, data),
    });

    $.export("$summary", "Successfully extracted text from document.");
    return response;
  },
};

Action Configuration#

This component may be configured based on the props defined in the component code. Pipedream automatically prompts for input values in the UI.

Label	Prop	Type	Description
LLMWhisperer	`app`	`app`	This component uses the LLMWhisperer app.
Processing Mode	`processingMode`	`string`	Select a value from the drop down menu:`ocrtext`
Output Mode	`outputMode`	`string`	Select a value from the drop down menu:`line-printertext`
Page Seperator	`pageSeperator`	`string`	The string to be used as a page separator. Eg: `<<<`
Force Text Processing	`forceTextProcessing`	`boolean`	If set to true, the document will be processed as text only. If set to false, the document will be processed based on LLMWhisperer's chosed stratergy.
Pages To Extract	`pagesToExtract`	`string`	Define which pages to extract. By default all pages are extracted. You can specify which pages to extract with this parameter. Example `1-5,7,21-` will extract pages 1,2,3,4,5,7,21,22,23,24... till the last page.
Timeout	`timeout`	`integer`	The time in seconds after which the request will automatically switch to async mode. If a timeout occurs then the API will return a 202 message along with `whisper-hash` which can be used later to check processing status and retrieve the text. Refer to the async operation documentation for more information
Store Metadata for Highlighting	`storeMetadataForHighlighting`	`boolean`	If set to true, metadata required for the highlighting is stored. If you do not require highlighting API, set this to false. Note that setting this to true will store your text in our servers
Median Filter Size	`medianFilterSize`	`integer`	The size of the median filter to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.
Gaussian Blur Radius	`gaussianBlurRadius`	`integer`	The radius of the gaussian blur to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.
OCR Provider	`ocrProvider`	`string`	Select a value from the drop down menu:`simpleadvanced`
Line Splitter Tolerance	`lineSplitterTolerance`	`string`	Factor to decide when to move text to the next line when it is above or below the baseline. The default value of `0.4` signifies 40% of the average character height.
Horizontal Stretch Factor	`horizontalStretchFactor`	`string`	Factor by which a horizontal stretch has to applied. It defaults to `1.0`. A stretch factor of `1.1` would mean at 10% stretch factor applied. Normally this factor need not be adjusted. You might want to use this parameter when multi column layouts back into each other. For example in a two column layout, the two columns get merged into one.
URL In Post	`urlInPost`	`boolean`	If set to `true`, the headers will be set to `text/plain`. If set to `false`, the headers will be set to `application/octet-stream`.

Extract Text with LLMWhisperer API on New Blog Posts from HubSpot API

Pipedream makes it easy to connect APIs for LLMWhisperer, HubSpot and 2,200+ other apps remarkably fast.

Trusted by 1,000,000+ developers from startups to Fortune 500 companies

Developers ♥ Pipedream

1-24of2,200+apps by most popular

1
-
24
of
2,200+
apps by most popular