diff --git a/docs/features/chat-conversations/audio/speech-to-text/azure-stt-integration.md b/docs/features/chat-conversations/audio/speech-to-text/azure-stt-integration.md new file mode 100644 index 0000000000..66562bbf30 --- /dev/null +++ b/docs/features/chat-conversations/audio/speech-to-text/azure-stt-integration.md @@ -0,0 +1,152 @@ +--- +sidebar_position: 2 +title: "Azure AI Speech STT" +--- + +# Using Azure AI Speech for Speech-to-Text + +This guide covers how to use Azure AI Speech for Speech-to-Text with Open WebUI. Open WebUI sends each recording to Azure's [fast transcription API](https://learn.microsoft.com/azure/ai-services/speech-service/fast-transcription-create), which returns the text in a single request. + +:::tip Looking for TTS? +See the companion guide: [Using Azure AI Speech for Text-to-Speech](/features/chat-conversations/audio/text-to-speech/azure-tts-integration) +::: + +## Requirements + +- An Azure Speech resource (Microsoft's documentation also calls it a Foundry resource for Speech) in a region that offers **fast transcription**. Not every region does: check the **Fast transcription** column of Microsoft's [Speech service regions table](https://learn.microsoft.com/azure/ai-services/speech-service/regions?tabs=stt). +- The resource's key, from **Resource Management** > **Keys and Endpoint** in the Azure portal. Either of the two keys works. +- Open WebUI installed and running + +:::caution Azure Government +Microsoft lists fast transcription as unsupported in Azure Government, so this engine cannot be used there. See [Speech service in sovereign clouds](https://learn.microsoft.com/azure/ai-services/speech-service/sovereign-clouds). +::: + +## Quick Setup (UI) + +1. Click your **profile icon** (bottom-left corner) +2. Select **Admin Panel**, then **Settings**. Settings opens in a window. +3. In the sidebar, choose **Audio** under **Admin**, in the *Experience* group. The other **Audio** entry, in the *Preferences* group, holds your personal audio settings. On a narrow screen these labels are hidden, so open the admin one directly at `/?settings=admin:audio` instead. +4. Configure the following: + +| Setting | Value | +|---------|-------| +| **Speech-to-Text Engine** | `Azure AI Speech` | +| **API Key** | Your Speech resource key | +| **Azure Region** | The region of your Speech resource, as an identifier such as `westus` or `westeurope`. Left blank, Open WebUI uses `eastus` | +| **Language Locales** | The languages your users speak, separated by commas with no spaces, for example `en-US,de-DE`. See [Language Locales](#language-locales) | +| **Endpoint URL** | Leave blank, unless you use your resource's own endpoint (see [Endpoint URL](#endpoint-url)) | +| **Max Speakers** | Leave blank to use the default of `3` | + +5. Click **Save** + +The key only works with the region the resource was created in. Microsoft notes that keys are region-scoped, so a key used with a different region fails to authenticate. + +## Language Locales + +Azure works out which language is being spoken from the list of locales Open WebUI sends with each recording: + +- When **Language Locales** is set, Azure uses those locales as its candidates. Microsoft notes that if none of them is in the audio, the service tries to identify the language itself. +- When it is left blank, Open WebUI sends this list: `en-US`, `es-ES`, `es-MX`, `fr-FR`, `hi-IN`, `it-IT`, `de-DE`, `en-GB`, `en-IN`, `ja-JP`, `ko-KR`, `pt-BR` and `zh-CN`. + +Azure identifies one main language per recording. Microsoft notes that a smaller, accurate set of candidate locales can improve detection, so list only the languages your users actually speak. For the locales Azure supports, see [Language and voice support](https://learn.microsoft.com/azure/ai-services/speech-service/language-support?tabs=stt). + +Separate locales with commas only. Open WebUI does not trim spaces, so `en-US, de-DE` sends `" de-DE"` with a leading space. + +The **Language** each user can set under **Settings** → **Audio** is not used by this engine. Only the admin's **Language Locales** list is sent to Azure. + +## Endpoint URL + +By default, Open WebUI sends requests to `https://.api.cognitive.microsoft.com`, built from **Azure Region**. **Endpoint URL** replaces that address, and Open WebUI adds `/speechtotext/transcriptions:transcribe?api-version=2024-11-15` to it. Enter only the scheme and host, with no path. + +You can use your resource's own endpoint from **Keys and Endpoint**, for example `https://.cognitiveservices.azure.com`. Microsoft's fast transcription examples call the same path on that address. Leave off the trailing `/` that Azure shows, because Open WebUI adds the path directly and would otherwise double the slash. + +## Max Speakers + +Open WebUI always asks Azure to separate speakers (diarization), with **Max Speakers** as the upper limit. The text Open WebUI keeps is the combined transcript, without speaker labels. + +## Environment Variables Setup + +If you prefer to configure via environment variables: + +```yaml +services: + open-webui: + image: ghcr.io/open-webui/open-webui:main + environment: + - AUDIO_STT_ENGINE=azure + - AUDIO_STT_AZURE_API_KEY=your-speech-resource-key + - AUDIO_STT_AZURE_REGION=westeurope + - AUDIO_STT_AZURE_LOCALES=en-US,de-DE + # ... other configuration +``` + +### All Azure STT Environment Variables + +| Variable | Description | Default | +|----------|-------------|---------| +| `AUDIO_STT_ENGINE` | Set to `azure` | empty (uses local Whisper) | +| `AUDIO_STT_AZURE_API_KEY` | Your Speech resource key | empty | +| `AUDIO_STT_AZURE_REGION` | Region identifier of your Speech resource | empty, which uses `eastus` | +| `AUDIO_STT_AZURE_LOCALES` | Comma-separated locales, no spaces | empty, which sends the 13 locales listed above | +| `AUDIO_STT_AZURE_BASE_URL` | Replaces `https://.api.cognitive.microsoft.com` | empty | +| `AUDIO_STT_AZURE_MAX_SPEAKERS` | Upper limit for speaker separation | empty, which uses `3` | + +:::info Settings saved in the UI take precedence +Saving the admin **Audio** settings stores every field on that tab, empty ones included, and choosing an engine saves too. With the default `ENABLE_PERSISTENT_CONFIG=true`, those stored values take precedence over the `AUDIO_*` environment variables from then on. See [`ENABLE_PERSISTENT_CONFIG`](/reference/env-configuration#enable_persistent_config). +::: + +## Long Recordings + +Before transcription, Open WebUI converts recordings in formats outside its supported list to MP3. Audio larger than 20 MB is re-encoded to a smaller MP3 and, if still too large, split into pieces of up to 20 MB. Each piece is sent to Azure as its own request, so the language is detected separately for each piece, and the texts are joined. + +The slim image, or `BYPASS_PYDUB_PREPROCESSING=true`, skips this step and sends files whole. Open WebUI then rejects files over 200 MB, and Microsoft's reference for the API version Open WebUI uses (`2024-11-15`) limits audio to under 2 hours. + +## Using STT + +1. Click the **microphone icon** in the chat input +2. Speak your message +3. Click the checkmark (**Confirm recording**) when you finish +4. Your speech is transcribed and appears in the message box, or is sent right away if **Instant Auto-Send After Voice Transcription** is on in your **Audio** settings + +## Troubleshooting + +### "Azure API key and region are required for Azure STT" + +No API key is saved. Enter your Speech resource key in **API Key** (or set `AUDIO_STT_AZURE_API_KEY`) and save. + +### "External: Access denied due to invalid subscription key or wrong API endpoint" + +Azure rejected the key. Check that: + +1. The key is copied correctly from **Keys and Endpoint** +2. **Azure Region** is the region your Speech resource is in +3. **Endpoint URL**, if set, is your resource's endpoint with no path + +### "External: Resource not found" + +Azure did not recognize the address. If **Endpoint URL** is set, make sure it has no path after the host. + +### Language could not be identified + +When Azure cannot identify a language, or finds several with none dominant, Open WebUI shows Azure's own error message. Set **Language Locales** to the languages actually spoken in your recordings. + +### Other errors + +Open WebUI shows Azure's own message when the audio is empty or too long, and for the two language identification errors. Other errors in Azure's transcription error format show "An error occurred during transcription.", and Azure's error code and message are written to the server log, in a line containing `Azure STT error`. + +A few other messages come from Open WebUI itself: + +- **A message starting with "Failed to parse Azure response:"**: Azure answered, but without any transcript text. +- **A message starting with "Error transcribing chunk:"**: the request never reached Azure, for example because the region identifier is mistyped. + +If the key and region are correct and transcription still fails, confirm that your region offers fast transcription. + +For more troubleshooting, see the [Audio Troubleshooting Guide](/troubleshooting/audio). + +## Cost Considerations + +Azure bills Speech usage under your Azure subscription. Check [Azure's Speech pricing page](https://azure.microsoft.com/pricing/details/cognitive-services/speech-services/) for current rates. Microsoft's [quotas and limits](https://learn.microsoft.com/azure/ai-services/speech-service/speech-services-quotas-and-limits) list fast transcription limits for the Standard (S0) tier only. + +:::tip +For free STT, use **Local Whisper** (the default) or the browser's **Web API** for basic transcription. +::: diff --git a/docs/features/chat-conversations/audio/speech-to-text/env-variables.md b/docs/features/chat-conversations/audio/speech-to-text/env-variables.md index f67e82dfcd..65a9ae3018 100644 --- a/docs/features/chat-conversations/audio/speech-to-text/env-variables.md +++ b/docs/features/chat-conversations/audio/speech-to-text/env-variables.md @@ -58,13 +58,15 @@ If using the `:cuda` Docker image with an older GPU, set `WHISPER_COMPUTE_TYPE=f ### Azure STT +See the [Azure AI Speech STT guide](/features/chat-conversations/audio/speech-to-text/azure-stt-integration) for setup. + | Variable | Description | Default | |----------|-------------|---------| | `AUDIO_STT_AZURE_API_KEY` | Azure Cognitive Services API key | empty | -| `AUDIO_STT_AZURE_REGION` | Azure region | `eastus` | -| `AUDIO_STT_AZURE_LOCALES` | Comma-separated locales (e.g., `en-US,de-DE`) | auto | +| `AUDIO_STT_AZURE_REGION` | Azure region | empty, which uses `eastus` | +| `AUDIO_STT_AZURE_LOCALES` | Comma-separated locales, no spaces (e.g., `en-US,de-DE`) | empty, which sends a built-in list of 13 locales | | `AUDIO_STT_AZURE_BASE_URL` | Custom Azure base URL (optional) | empty | -| `AUDIO_STT_AZURE_MAX_SPEAKERS` | Max speakers for diarization | `3` | +| `AUDIO_STT_AZURE_MAX_SPEAKERS` | Max speakers for diarization | empty, which uses `3` | ### Deepgram STT @@ -113,9 +115,11 @@ When `AUDIO_TTS_ENGINE=mistral`, Open WebUI uses `mistral-tts-latest` when `AUDI ### Azure TTS +See the [Azure AI Speech TTS guide](/features/chat-conversations/audio/text-to-speech/azure-tts-integration) for setup. + | Variable | Description | Default | |----------|-------------|---------| -| `AUDIO_TTS_AZURE_SPEECH_REGION` | Azure Speech region | `eastus` | +| `AUDIO_TTS_AZURE_SPEECH_REGION` | Azure Speech region | empty (speech uses `eastus`; without a base URL, the voice list does not load) | | `AUDIO_TTS_AZURE_SPEECH_BASE_URL` | Custom Azure Speech base URL (optional) | empty | | `AUDIO_TTS_AZURE_SPEECH_OUTPUT_FORMAT` | Audio output format | `audio-24khz-160kbitrate-mono-mp3` | diff --git a/docs/features/chat-conversations/audio/speech-to-text/stt-config.md b/docs/features/chat-conversations/audio/speech-to-text/stt-config.md index 82edf8cfb1..41ab0e71c6 100644 --- a/docs/features/chat-conversations/audio/speech-to-text/stt-config.md +++ b/docs/features/chat-conversations/audio/speech-to-text/stt-config.md @@ -22,7 +22,7 @@ The following speech-to-text providers are supported: | Self-hosted (OpenAI-compatible) | ✅ | [Self-Hosted STT Guide](/features/chat-conversations/audio/speech-to-text/self-hosted-stt) (uses an authenticated server) | | Mistral (Voxtral) | ✅ | [Mistral Voxtral Guide](/features/chat-conversations/audio/speech-to-text/mistral-voxtral-integration) | | Deepgram | ✅ | N/A | -| Azure | ✅ | N/A | +| Azure | ✅ | [Azure AI Speech STT Guide](/features/chat-conversations/audio/speech-to-text/azure-stt-integration) | **Web API** provides STT via the browser's built-in speech recognition (no API key needed, configured in user settings). diff --git a/docs/features/chat-conversations/audio/text-to-speech/azure-tts-integration.md b/docs/features/chat-conversations/audio/text-to-speech/azure-tts-integration.md new file mode 100644 index 0000000000..c801d27bb3 --- /dev/null +++ b/docs/features/chat-conversations/audio/text-to-speech/azure-tts-integration.md @@ -0,0 +1,145 @@ +--- +sidebar_position: 2 +title: "Azure AI Speech TTS" +--- + +# Using Azure AI Speech for Text-to-Speech + +This guide covers how to use Azure AI Speech for Text-to-Speech with Open WebUI. Open WebUI sends text to Azure's [text to speech REST API](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech) and plays back the audio it returns. + +:::tip Looking for STT? +See the companion guide: [Using Azure AI Speech for Speech-to-Text](/features/chat-conversations/audio/speech-to-text/azure-stt-integration) +::: + +## Requirements + +- An Azure Speech resource (Microsoft's documentation also calls it a Foundry resource for Speech). See Microsoft's [Speech service regions table](https://learn.microsoft.com/azure/ai-services/speech-service/regions?tabs=tts) for text to speech availability by region. +- The resource's key, from **Resource Management** > **Keys and Endpoint** in the Azure portal. Either of the two keys works. +- Open WebUI installed and running + +## Quick Setup (UI) + +1. Click your **profile icon** (bottom-left corner) +2. Select **Admin Panel**, then **Settings**. Settings opens in a window. +3. In the sidebar, choose **Audio** under **Admin**, in the *Experience* group. The other **Audio** entry, in the *Preferences* group, holds your personal audio settings. On a narrow screen these labels are hidden, so open the admin one directly at `/?settings=admin:audio` instead. +4. Configure the following: + +| Setting | Value | +|---------|-------| +| **Text-to-Speech Engine** | `Azure AI Speech` | +| **API Key** | Your Speech resource key | +| **Azure Region** | The region of your Speech resource, as an identifier such as `westus` or `westeurope` | +| **Endpoint URL** | Leave blank, unless you use a sovereign cloud (see [Endpoint URL](#endpoint-url)) | + +5. Click **Save**. Choosing the engine already saves on its own and clears **TTS Voice**, which is why the voice comes after this step. +6. Switch to another settings tab and back to **Audio**. The voice list is fetched when the Audio tab opens, so it only loads with your key after this. +7. Configure the voice and format: + +| Setting | Value | +|---------|-------| +| **TTS Voice** | Type part of a voice name or language to search, then pick a voice, for example `en-US-JennyNeural` | +| **Output format** | Leave the default `audio-24khz-160kbitrate-mono-mp3` (see [Output Format](#output-format)) | + +8. Click **Save** + +The key only works with the region the resource was created in. Microsoft notes that keys are region-scoped, so a key used with a different region fails to authenticate. + +:::caution Always set the region +If **Azure Region** and **Endpoint URL** are both blank, speech is still requested from `eastus`, but the voice list is not loaded at all. Set the region your resource is in. +::: + +## Choosing Voices + +Open WebUI loads the voice list for your region from Azure. Each row in **TTS Voice** shows the voice's full name, such as `en-US-JennyNeural`, with its display label next to it in grey. Only a few rows show at a time, so type part of a name or a language, such as `en-GB`, to narrow the list. + +The full name is what Open WebUI sends to Azure, and it takes the speech language from the first two parts of that name (`en-US` from `en-US-JennyNeural`). If you type a voice instead of picking one, use the full name. + +Which voice is used for a reply: + +1. The model's own **TTS Voice**, if one is set when editing the model in **Workspace** → **Models** +2. Otherwise, the voice the user picked under **Set Voice** in their personal **Audio** settings, as long as the admin's **TTS Voice** is still the one that was set when they picked it. If the admin changes it, users hear the new default until they pick again. +3. Otherwise, the admin's **TTS Voice** + +For the full list of voices and languages, see Microsoft's [Language and voice support](https://learn.microsoft.com/azure/ai-services/speech-service/language-support?tabs=tts). + +## Endpoint URL + +By default, Open WebUI sends requests to `https://.tts.speech.microsoft.com`, built from **Azure Region**. **Endpoint URL** replaces that address. Open WebUI adds `/cognitiveservices/v1` to it to request speech and `/cognitiveservices/voices/list` to load the voice list, so enter only the scheme and host, with no path. + +This is not your resource's endpoint from **Keys and Endpoint**, the one the [STT guide](/features/chat-conversations/audio/speech-to-text/azure-stt-integration#endpoint-url) can use. Microsoft documents the voice list at `/tts/cognitiveservices/voices/list` on that address, while Open WebUI requests `/cognitiveservices/voices/list`, so the voice list would not load. + +For the sovereign clouds, Microsoft documents these addresses, with `` replaced by your region identifier: + +| Cloud | Endpoint URL | +|-------|--------------| +| Azure Government | `https://.tts.speech.azure.us` | +| Azure operated by 21Vianet | `https://.tts.speech.azure.cn` | + +See [Speech service in sovereign clouds](https://learn.microsoft.com/azure/ai-services/speech-service/sovereign-clouds) for the region identifiers. + +## Output Format + +**Output format** is sent to Azure as the audio format to return. Keep one of the MP3 formats, such as the default `audio-24khz-160kbitrate-mono-mp3`. Open WebUI stores the audio it receives as an `.mp3` file and, except in the slim image, serves it as MP3 whatever format Azure returned. The slim image refuses audio types outside its list of browser-playable formats, such as MP3, WAV and Ogg. Microsoft's [audio outputs list](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech#audio-outputs) has every value. + +## Environment Variables Setup + +If you prefer to configure via environment variables: + +```yaml +services: + open-webui: + image: ghcr.io/open-webui/open-webui:main + environment: + - AUDIO_TTS_ENGINE=azure + - AUDIO_TTS_API_KEY=your-speech-resource-key + - AUDIO_TTS_AZURE_SPEECH_REGION=westeurope + - AUDIO_TTS_VOICE=en-US-JennyNeural + # ... other configuration +``` + +Set `AUDIO_TTS_VOICE` to an Azure voice name. Its default, `alloy`, is an OpenAI voice. + +### All Azure TTS Environment Variables + +| Variable | Description | Default | +|----------|-------------|---------| +| `AUDIO_TTS_ENGINE` | Set to `azure` | empty (uses browser-only TTS) | +| `AUDIO_TTS_API_KEY` | Your Speech resource key (shared with the ElevenLabs engine) | empty | +| `AUDIO_TTS_AZURE_SPEECH_REGION` | Region identifier of your Speech resource | empty (speech uses `eastus`; without a base URL, the voice list does not load) | +| `AUDIO_TTS_AZURE_SPEECH_BASE_URL` | Replaces `https://.tts.speech.microsoft.com` | empty | +| `AUDIO_TTS_AZURE_SPEECH_OUTPUT_FORMAT` | Audio output format | `audio-24khz-160kbitrate-mono-mp3` | +| `AUDIO_TTS_VOICE` | Full Azure voice name | `alloy` (not an Azure voice) | + +:::info Settings saved in the UI take precedence +Saving the admin **Audio** settings stores every field on that tab, empty ones included, and choosing an engine saves too. With the default `ENABLE_PERSISTENT_CONFIG=true`, those stored values take precedence over the `AUDIO_*` environment variables from then on. See [`ENABLE_PERSISTENT_CONFIG`](/reference/env-configuration#enable_persistent_config). +::: + +## Testing TTS + +1. Start a new chat +2. Send a message to any model +3. Click the **speaker icon** on the AI response to hear it read aloud + +## Troubleshooting + +### No voices shown in the list + +1. Confirm **API Key** and **Azure Region** are set and saved +2. Switch to another settings tab and back to **Audio** to reload the list +3. Check the Open WebUI logs for `Error fetching Azure voices` + +### Speech fails with a 401 error + +Azure rejected the key. Check that the key is copied correctly from **Keys and Endpoint** and that **Azure Region** is the region your Speech resource is in. + +### Speech fails with another error + +1. Confirm **TTS Voice** is set to a full Azure voice name from the list +2. Confirm **Output format** is one of the MP3 values from Microsoft's [audio outputs list](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech#audio-outputs) +3. Check the Open WebUI logs for the full error + +For broader audio debugging, see the [Audio Troubleshooting Guide](/troubleshooting/audio). + +## Cost Considerations + +Azure bills Speech usage under your Azure subscription. Check [Azure's Speech pricing page](https://azure.microsoft.com/pricing/details/cognitive-services/speech-services/) for current rates.