docs: add Azure AI Speech setup guides for speech-to-text and text-to-speech - #1449
Open
silentoplayz wants to merge 1 commit into
Open
silentoplayz wants to merge 1 commit into
silentoplayz wants to merge 1 commit into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Azure AI Speech is one of the speech-to-text and text-to-speech engines in Admin → Audio, but the docs only listed its environment variables, and the speech-to-text provider table said "N/A" for its guide. This adds a setup guide for each, in the same layout as the Mistral guides:
speech-to-text/azure-stt-integration.md): the region requirement for fast transcription, Quick Setup, how Language Locales works (including the 13 locales sent when it is blank, and that spaces are not trimmed), Endpoint URL, Max Speakers, environment variables, how long recordings are split before they reach Azure, and the error messages users see.text-to-speech/azure-tts-integration.md): Quick Setup, why the voice list needs the region, how voices are picked and which one wins, Endpoint URL including the sovereign cloud addresses, keeping an MP3 Output format, environment variables, and troubleshooting.It also links the Azure row in the speech-to-text provider table and the Azure sections of the environment variables page to the new guides, and corrects four Azure defaults on that page to what the code does when the variable is empty.
Related issue or discussion
None. Azure Speech setup was listed as missing in #2 (comment) ("Azure Speech Service Integration").
Checklist
Notes for reviewers
I did not test this against a live Azure account. Every statement was instead checked against these sources:
devat 015dbc8:_transcribe_azure,_tts_azure, the Azure branch ofget_available_voices,transcribeand the audio preprocessing inrouters/audio.py, the admin and user Audio settings, the voice picker, and theen-USlabels.2024-11-15Transcribe reference (the API version Open WebUI sends), the text to speech REST API, the regions table, sovereign clouds, and quotas and limits./speechtotext/transcriptions:transcribe?api-version=2024-11-15,/cognitiveservices/v1,/cognitiveservices/voices/liston the regional hosts) returns401, where a wrong path or API version returns404. The two quoted speech-to-text error messages are Azure's actual401and404responses.Two things the guides state that are easy to get wrong:
eastus, but the voice list does not load.The pages built and rendered with
docusaurus start, and every internal link and anchor resolves.