Speech to Text (Whisper)
Audience:
Low-code EngineersSkill Prerequisites:
Tokens,Connectors
Sends an MP3 file from your site to OpenAI's Whisper model, and saves the transcript in a token. The transcript can be plain text, JSON, or subtitles in SRT or VTT format.
This action is part of the AI add-on (PlantAnApp.OpenAi). The add-on is installed separately and needs the AI feature in your license. If it isn't licensed, the action fails with an "AI package is unlicensed" error. If you don't see the OpenAi actions, the add-on isn't installed.
Typical Use Cases
- Transcribe a voice message or recorded call uploaded through a form
- Create SRT or VTT subtitles for a video's audio track
- Turn a meeting recording into text, then summarize it with Summarize Content
Don't use it to
- Transcribe files that aren't MP3. Convert them to MP3 first.
- Transcribe audio from a URL on another site. The file has to be in your site's file system. Use Save File to download it first.
- Transcribe with Azure OpenAI. See Considerations.
- Process recordings you're not allowed to send to a third party.
Related Actions
| Action Name | Description |
|---|---|
| Summarize Content | Summarizes the transcript. |
| Create Articles | Turns a long transcript into separate articles. |
| Chat | Answers questions about the transcript, or reformats it. |
| Save File | Downloads an audio file, or saves a subtitle transcript as a file. |
| Test Connector | Checks that the OpenAI connector works. |
Input Parameter Reference
| Parameter | Description | Supports Tokens | Default | Required |
|---|---|---|---|---|
| Provider | The AI service: OpenAI or Azure. Only OpenAI works in 1.28. | No | OpenAI | Yes |
| OpenAI Connector | OpenAI only. The OpenAI connector that holds your API key. See Connectors. | No | none selected | Yes |
| Azure Connector | Azure only. A Microsoft Azure OpenAI connector. | No | none selected | Yes |
| Audio file | The MP3 file to transcribe. It can be a file ID, a relative URL, an absolute URL, a LinkClick URL or a physical path, as long as the file is in your site's file system, for example [RecordingFileId] or /Portals/0/Recordings/call.mp3. | Yes | empty string | Yes |
| Response Format | The format of the transcript. See Response formats. | Yes | none selected | No |
| Prompt | Optional text that guides the transcript, for example product names, spellings, or a sample of the style you want. | Yes | empty string | No |
Output Parameters Reference
| Parameter | Description |
|---|---|
| Output Token Name | Token that receives the transcript exactly as OpenAI returns it, for example Transcript. It's required. If it's empty, the action fails with "There is no output token defined to save the transcript into." |
Response formats
| Response Format | What the token receives |
|---|---|
json | A JSON object with the transcript in a text property, for example {"text":"Hello, thanks for calling."}. OpenAI uses this format when Response Format is empty. |
text | The transcript as plain text. |
srt | Subtitles in SRT format, with numbered, timed segments. |
verbose_json | A JSON object with the transcript, the detected language, the duration and timed segments. |
vtt | Subtitles in WebVTT format. |
To use a value from json or verbose_json, parse the token with Parse JSON Into Tokens. Choose text if you only need the words.
Considerations
- Your audio is sent to OpenAI. The whole file leaves your server and is processed by OpenAI under your account's terms. Get consent before sending recordings of people, and don't send confidential audio unless your agreement with OpenAI allows it.
- Only MP3 files. The action checks the file extension. Any other extension fails with "File format ... is not supported for speech to text."
- File size. OpenAI limits the size of audio uploads (25 MB at the time of writing). The action doesn't check the size first, so a bigger file fails with OpenAI's error. Split or compress long recordings.
- Missing files. If the file isn't found, the action fails with "File '...' not found."
- The model is fixed. The action always uses OpenAI's
whisper-1model. There's no language setting. Whisper detects the language, and you can hint at it inPrompt. - Azure doesn't work in 1.28. The action has no Deployment Id for Azure, and it sends the key in a way Azure doesn't accept. Use the OpenAI provider.
- No error handling of its own. The action doesn't have
On ErrororIgnore Errors. If OpenAI returns an error, the action fails with "OpenAi Whisper returned error" followed by the status and OpenAI's message. Use theOn Erroractions of Execute Actions to handle it. - Long runs. The request waits for the whole transcript. Long recordings can take a while, so consider running it in a workflow instead of a form submit.
- Usage isn't tracked. Unlike the other AI actions, this action doesn't create an AI usage record.
Examples
To understand how to use the below examples, please see Running Examples.
After importing the example, select your OpenAI connector in the action. The connector ID in the JSON is a placeholder.
1. Transcribe an uploaded voice message
This action transcribes the MP3 file whose ID is in the VoiceMessageFileId token, and saves the plain text in the Transcript token. The prompt helps Whisper spell product names correctly.
{
"Title": "Speech to Text (Whisper)",
"ActionType": "OpenAi.Whisper",
"Description": "Transcribe the voice message",
"Condition": "\"[VoiceMessageFileId]\" != \"\"",
"Parameters": {
"Provider": "openAi",
"OpenAiConnector": {
"Entry": "00000000-0000-0000-0000-000000000000"
},
"AudioFileIdentifier": "[VoiceMessageFileId]",
"ResponseFormat": {
"Expression": "",
"Value": "text",
"IsExpression": false,
"Parameters": {}
},
"Prompt": "A customer calling Plant an App support about workflows and connectors.",
"OutputTokenName": "Transcript"
}
}
Revised 09/27/2026