Extract Text with LLMWhisperer API

Pipedream makes it easy to connect APIs for LLMWhisperer and 2,700+ other apps remarkably fast.

Trigger workflow on

HTTP requests, schedules and app events

Next, do this

Extract Text with the LLMWhisperer API

No credit card required

▶

Watch us build a workflow

8 min

Watch now ➜

Trusted by 1,000,000+ developers from startups to Fortune 500 companies

Developers ♥ Pipedream

Chase Roberts
@chsrbrts
@benedictevans If you haven’t used @pipedream yet, then you haven’t lived.
✨Ellie Day✨
@heyellieday
Evaluation update: @pipedream has quite literally been a dream to work with! I’m excited to leverage this tool for all the various workflows I need to write. I’m currently at 11k invocations a day from the initial workflows I’ve written in the past couple weeks.
Michael Braedley
@MBraedley
Update: I got it working properly, and it's working so well that I'm dropping IFTTT. @pipedream can do everything that IFTTT basic can, and most (if not all things) IFTTT pro can for free or at a reasonable price if you need it. I am recommending it for basically any power user.
Thomas Cutting
@mrthomascutting
Want quick+dirty integrations for a serverless workflow - @pipedream is my new go-to 😃
Matthew Roberts
@mattdotroberts
day 013 - finally hit node js. This is the secret sauce of taking #nocode projects that one step further. Pumped about getting deeper into @pipedream now
Kenneth Auchenberg 💭
@auchenberg
Yahoo Pipes is back! Kinda 😍 @pipedream
Raymond Camden 🥑
@raymondcamden
Awesome video by the @pipedream folks showing real time twitter sentiment analysis integrated with Google Sheets. This is where Pipedream *really* shines, connecting systems together in easy workflows.
Nacho Caballero
@nachocaballero
I couldn't recommend @pipedream more. It's an amazing service to integrate different APIs. Much more powerful than Zapier and more user-friendly than AWS Lambda. I'm very proud to wear this t-shirt #NoCode
Jason Snow
@jyksnw
Developed a working prototype environmental sensor IoT solution with @particle Photon, @pipedream, and @MongoDB with full graphing and alerting in less than a day! All amazing technology, will def. be exploring these more.
Steven Terrana
@steven_terrana
@burgwyn you've inspired me to finally set up my own blog. I'll make sure my first blog post explains the tech behind the setup. think @obsdmd + @GatsbyJS + @pipedream.
🚄 James Augeri, PhD
@DotDotJames
Want to low-code your back end, need more horsepower than @Bubble / @KnackHQ, or just miss Yahoo! Pipes? Check out @PipeDream
Sébastien Chopin
@Atinux
GitHub issues should be like @linear_app for maintainers. Looking forward more integrations with GH actions or tools like @pipedream 👀
Raul
@raul_predescu
If you're a dev and not using @pipedream, you're missing out. Been using it for months, daily. FREE for devs. Plenty of integrations and good limits. Absolutely love it.
Bruno Skvorc
@bitfalls
So @pipedream is pretty amazing. In 3 minutes I just made a flow which adds @rickastley's Never Gonna Give You Up to my @spotify playlist whenever a new pull request arrives in an old repo of mine.
Zach Lanich
@ZachLanich
Um, wow 🤯 @pipedream
Steven Bell
@bellontech
I just used @pipedream to build a Shopify App. Wow, they make small backed tasks easy.
Jay Hack 🎩🇺🇸
@_jayhack_
Very impressed with this bad boi - it reminds me of a @PalantirTech internal tool, but geared towards integrations instead of data analysis and far more customizable. Great expectations here 🚀🤩
Tree Sturgeon 🔥🚴‍♂️🌳
@philsturgeon
For context this is day 2 of a really challenging and stupid migration from Notion to @airtable with disparate/missing data. It's going better than expected and thanks to @pipedream I don't have to bother the iOS dev to add W3W.

Getting Started#

Create a workflow to Extract Text with the LLMWhisperer API. When you configure and deploy the workflow, it will run on Pipedream's servers 24x7 for free.

Configure the Extract Text action
1. Connect your LLMWhisperer account
2. Select a Processing Mode
3. Select a Output Mode
4. Optional- Configure Page Seperator
5. Optional- Configure Force Text Processing
6. Optional- Configure Pages To Extract
7. Optional- Configure Timeout
8. Optional- Configure Store Metadata for Highlighting
9. Optional- Configure Median Filter Size
10. Optional- Configure Gaussian Blur Radius
11. Optional- Select a OCR Provider
12. Optional- Configure Line Splitter Tolerance
13. Optional- Configure Horizontal Stretch Factor
14. Configure URL In Post
Select a trigger to run your workflow on HTTP requests, schedules or app events
Deploy the workflow
Send a test event to validate your setup
Turn on the trigger

Integrations#

Extract Text with LLMWhisperer API on New Requests (Payload Only) from HTTP / Webhook API

HTTP / Webhook + LLMWhisperer

Try it

Extract Text with LLMWhisperer API on New Submission from Typeform API

Typeform + LLMWhisperer

Try it

Extract Text with LLMWhisperer API on New Submission (Instant) from Jotform API

Jotform + LLMWhisperer

Try it

Extract Text with LLMWhisperer API on New Scheduled Tasks from Pipedream API

Pipedream + LLMWhisperer

Try it

Extract Text with LLMWhisperer API on New Download Counts from npm API

npm + LLMWhisperer

Try it

Details#

This is a pre-built, source-available component from Pipedream's GitHub repo. The component is developed by Pipedream and the community, and verified and maintained by Pipedream.

To contribute an update to an existing component or create a new component, create a PR on GitHub. If you're new to Pipedream component development, you can start with quickstarts for trigger span and action development, and then review the component API reference.

Extract Text on LLMWhisperer

Description:Convert your PDF/scanned documents to text format which can be used by LLMs. [See the documentation](https://docs.unstract.com/llm_whisperer/apis/llm_whisperer_text_extraction_api)

Version:0.0.1

Key:llmwhisperer-extract-text

View on GitHub

Code#

import fs from "fs";
import app from "../../llmwhisperer.app.mjs";

export default {
  key: "llmwhisperer-extract-text",
  name: "Extract Text",
  description: "Convert your PDF/scanned documents to text format which can be used by LLMs. [See the documentation](https://docs.unstract.com/llm_whisperer/apis/llm_whisperer_text_extraction_api)",
  version: "0.0.1",
  type: "action",
  props: {
    app,
    processingMode: {
      type: "string",
      label: "Processing Mode",
      description: "The processing mode to be used. Choose between `ocr` and `text`.",
      options: [
        "ocr",
        "text",
      ],
    },
    outputMode: {
      type: "string",
      label: "Output Mode",
      description: "The output mode to be used. Choose between `line-printer` and `text`.",
      options: [
        "line-printer",
        "text",
      ],
    },
    pageSeperator: {
      type: "string",
      label: "Page Seperator",
      description: "The string to be used as a page separator. Eg: `<<<`",
      optional: true,
    },
    forceTextProcessing: {
      type: "boolean",
      label: "Force Text Processing",
      description: "If set to true, the document will be processed as text only. If set to false, the document will be processed based on LLMWhisperer's chosed stratergy.",
      optional: true,
    },
    pagesToExtract: {
      type: "string",
      label: "Pages To Extract",
      description: "Define which pages to extract. By default all pages are extracted. You can specify which pages to extract with this parameter. Example `1-5,7,21-` will extract pages **1,2,3,4,5,7,21,22,23,24...** till the last page.",
      optional: true,
    },
    timeout: {
      type: "integer",
      label: "Timeout",
      description: "The time in seconds after which the request will automatically switch to async mode. If a timeout occurs then the API will return a 202 message along with `whisper-hash` which can be used later to check processing status and retrieve the text. Refer to the async operation documentation for more information",
      optional: true,
    },
    storeMetadataForHighlighting: {
      type: "boolean",
      label: "Store Metadata for Highlighting",
      description: "If set to true, metadata required for the highlighting is stored. If you do not require highlighting API, set this to false. Note that setting this to true will store your text in our servers",
      optional: true,
    },
    medianFilterSize: {
      type: "integer",
      label: "Median Filter Size",
      description: "The size of the median filter to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
    },
    gaussianBlurRadius: {
      type: "integer",
      label: "Gaussian Blur Radius",
      description: "The radius of the gaussian blur to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
    },
    ocrProvider: {
      type: "string",
      label: "OCR Provider",
      description: "The OCR provider to be used. Choose between `simple` and `advanced`. This parameter works only in on-prem version of LLMWhisperer.",
      optional: true,
      options: [
        "simple",
        "advanced",
      ],
    },
    lineSplitterTolerance: {
      type: "string",
      label: "Line Splitter Tolerance",
      description: "Factor to decide when to move text to the next line when it is above or below the baseline. The default value of `0.4` signifies 40% of the average character height.",
      optional: true,
    },
    horizontalStretchFactor: {
      type: "string",
      label: "Horizontal Stretch Factor",
      description: "Factor by which a horizontal stretch has to applied. It defaults to `1.0`. A stretch factor of `1.1` would mean at 10% stretch factor applied. Normally this factor need not be adjusted. You might want to use this parameter when multi column layouts back into each other. For example in a two column layout, the two columns get merged into one.",
      optional: true,
    },
    urlInPost: {
      type: "boolean",
      label: "URL In Post",
      description: "If set to `true`, the headers will be set to `text/plain`. If set to `false`, the headers will be set to `application/octet-stream`.",
      reloadProps: true,
      default: true,
    },
  },
  additionalProps() {
    const { urlInPost } = this;
    return {
      data: {
        type: "string",
        label: urlInPost
          ? "Document URL"
          : "Document Path",
        description: urlInPost
          ? "The URL of the document to process."
          : "Document path of the file previously downloaded in Pipedream E.g. (`/tmp/my-file.txt`). [Download a file to the `/tmp` directory](https://pipedream.com/docs/code/nodejs/http-requests/#download-a-file-to-the-tmp-directory)",
      },
    };
  },
  methods: {
    getHeaders(urlInPost) {
      return {
        "Content-Type": urlInPost
          ? "text/plain"
          : "application/octet-stream",
      };
    },
    getData(urlInPost, data) {
      return urlInPost
        ? data
        : fs.readFileSync(data);
    },
    extractText(args = {}) {
      return this.app.post({
        path: "/whisper",
        ...args,
      });
    },
  },
  async run({ $ }) {
    const {
      extractText,
      getHeaders,
      getData,
      urlInPost,
      processingMode,
      outputMode,
      pageSeperator,
      forceTextProcessing,
      pagesToExtract,
      timeout,
      storeMetadataForHighlighting,
      medianFilterSize,
      gaussianBlurRadius,
      ocrProvider,
      lineSplitterTolerance,
      horizontalStretchFactor,
      data,
    } = this;

    const response = await extractText({
      $,
      headers: getHeaders(urlInPost),
      params: {
        url_in_post: urlInPost,
        processing_mode: processingMode,
        output_mode: outputMode,
        page_seperator: pageSeperator,
        force_text_processing: forceTextProcessing,
        pages_to_extract: pagesToExtract,
        timeout,
        store_metadata_for_highlighting: storeMetadataForHighlighting,
        median_filter_size: medianFilterSize,
        gaussian_blur_radius: gaussianBlurRadius,
        ocr_provider: ocrProvider,
        line_splitter_tolerance: lineSplitterTolerance,
        horizontal_stretch_factor: horizontalStretchFactor,
      },
      data: getData(urlInPost, data),
    });

    $.export("$summary", "Successfully extracted text from document.");
    return response;
  },
};

Configuration#

This component may be configured based on the props defined in the component code. Pipedream automatically prompts for input values in the UI and CLI.

Label	Prop	Type	Description
LLMWhisperer	`app`	`app`	This component uses the LLMWhisperer app.
Processing Mode	`processingMode`	`string`	Select a value from the drop down menu:`ocrtext`
Output Mode	`outputMode`	`string`	Select a value from the drop down menu:`line-printertext`
Page Seperator	`pageSeperator`	`string`	The string to be used as a page separator. Eg: `<<<`
Force Text Processing	`forceTextProcessing`	`boolean`	If set to true, the document will be processed as text only. If set to false, the document will be processed based on LLMWhisperer's chosed stratergy.
Pages To Extract	`pagesToExtract`	`string`	Define which pages to extract. By default all pages are extracted. You can specify which pages to extract with this parameter. Example `1-5,7,21-` will extract pages 1,2,3,4,5,7,21,22,23,24... till the last page.
Timeout	`timeout`	`integer`	The time in seconds after which the request will automatically switch to async mode. If a timeout occurs then the API will return a 202 message along with `whisper-hash` which can be used later to check processing status and retrieve the text. Refer to the async operation documentation for more information
Store Metadata for Highlighting	`storeMetadataForHighlighting`	`boolean`	If set to true, metadata required for the highlighting is stored. If you do not require highlighting API, set this to false. Note that setting this to true will store your text in our servers
Median Filter Size	`medianFilterSize`	`integer`	The size of the median filter to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.
Gaussian Blur Radius	`gaussianBlurRadius`	`integer`	The radius of the gaussian blur to be applied to the image. This is used to remove noise from the image. This parameter works only in on-prem version of LLMWhisperer.
OCR Provider	`ocrProvider`	`string`	Select a value from the drop down menu:`simpleadvanced`
Line Splitter Tolerance	`lineSplitterTolerance`	`string`	Factor to decide when to move text to the next line when it is above or below the baseline. The default value of `0.4` signifies 40% of the average character height.
Horizontal Stretch Factor	`horizontalStretchFactor`	`string`	Factor by which a horizontal stretch has to applied. It defaults to `1.0`. A stretch factor of `1.1` would mean at 10% stretch factor applied. Normally this factor need not be adjusted. You might want to use this parameter when multi column layouts back into each other. For example in a two column layout, the two columns get merged into one.
URL In Post	`urlInPost`	`boolean`	If set to `true`, the headers will be set to `text/plain`. If set to `false`, the headers will be set to `application/octet-stream`.

Try it

Highlight Locations with the LLMWhisperer API

Generate highlight locations for a search term in the document. See the documentation

Try it

Retrieve Extracted Text with the LLMWhisperer API

Retrieve the extracted text executed through the whisper API. This can be used to retrieve the text of the conversion process when the conversion is done in async mode. See the documentation

Try it