Skip to content

docs: add Azure AI Speech setup guides for speech-to-text and text-to-speech - #1449

Open
silentoplayz wants to merge 1 commit into
open-webui:mainfrom
silentoplayz:docs/azure-speech-guides
Open

silentoplayz wants to merge 1 commit into
open-webui:mainfrom
silentoplayz:docs/azure-speech-guides

Conversation

@silentoplayz

Copy link
Copy Markdown
Collaborator

Summary

Azure AI Speech is one of the speech-to-text and text-to-speech engines in Admin → Audio, but the docs only listed its environment variables, and the speech-to-text provider table said "N/A" for its guide. This adds a setup guide for each, in the same layout as the Mistral guides:

  • Speech-to-text (speech-to-text/azure-stt-integration.md): the region requirement for fast transcription, Quick Setup, how Language Locales works (including the 13 locales sent when it is blank, and that spaces are not trimmed), Endpoint URL, Max Speakers, environment variables, how long recordings are split before they reach Azure, and the error messages users see.
  • Text-to-speech (text-to-speech/azure-tts-integration.md): Quick Setup, why the voice list needs the region, how voices are picked and which one wins, Endpoint URL including the sovereign cloud addresses, keeping an MP3 Output format, environment variables, and troubleshooting.

It also links the Azure row in the speech-to-text provider table and the Azure sections of the environment variables page to the new guides, and corrects four Azure defaults on that page to what the code does when the variable is empty.

Related issue or discussion

None. Azure Speech setup was listed as missing in #2 (comment) ("Azure Speech Service Integration").

Checklist

  • I have reviewed the relevant documentation and matched the existing style.
  • This PR meets Open WebUI's contribution standards: it is accurate, relevant to users, narrowly scoped, maintainable, and not promotional content, advertising, lead generation, SEO placement, or a request to list a product, service, provider, integration, gateway, tool, or company primarily for visibility.
  • I understand that PRs that do not meet these standards may be closed without review and will not be merged. Repeated, low-quality, off-topic, promotional, or intentionally misleading submissions may result in the contributor being blocked from future participation in Open WebUI repositories.

Notes for reviewers

I did not test this against a live Azure account. Every statement was instead checked against these sources:

  • Open WebUI on dev at 015dbc8: _transcribe_azure, _tts_azure, the Azure branch of get_available_voices, transcribe and the audio preprocessing in routers/audio.py, the admin and user Audio settings, the voice picker, and the en-US labels.
  • Microsoft's documentation: the fast transcription guide and the 2024-11-15 Transcribe reference (the API version Open WebUI sends), the text to speech REST API, the regions table, sovereign clouds, and quotas and limits.
  • The live endpoints, without a key: each URL Open WebUI builds (/speechtotext/transcriptions:transcribe?api-version=2024-11-15, /cognitiveservices/v1, /cognitiveservices/voices/list on the regional hosts) returns 401, where a wrong path or API version returns 404. The two quoted speech-to-text error messages are Azure's actual 401 and 404 responses.

Two things the guides state that are easy to get wrong:

  • Microsoft lists fast transcription as unsupported in Azure Government, so the speech-to-text guide says that engine can't be used there.
  • With Azure Region blank, text-to-speech still speaks through eastus, but the voice list does not load.

The pages built and rendered with docusaurus start, and every internal link and anchor resolves.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant