diff --git a/content/docs/configuration/stt_tts.mdx b/content/docs/configuration/stt_tts.mdx index a6ea4336d..676922a01 100644 --- a/content/docs/configuration/stt_tts.mdx +++ b/content/docs/configuration/stt_tts.mdx @@ -182,15 +182,85 @@ You can specify either the full domain or just the instance name. If you provide Refer to the OpenAI Whisper section, adjusting the `url` and `model` as needed. -example - ```yaml speech: stt: openai: url: 'http://host.docker.internal:8080/v1/audio/transcriptions' model: 'whisper' - ``` +``` + +- #### FunASR with SenseVoice + +[FunASR](https://github.com/modelscope/FunASR) provides an OpenAI-compatible transcription endpoint. The following example uses the lightweight SenseVoice model for multilingual speech recognition, including Mandarin, Cantonese, English, Japanese, and Korean. + +Create a separate Python environment for the FunASR service: + +```bash +python -m venv .venv +source .venv/bin/activate +``` + +Install matching PyTorch and torchaudio builds for your platform using the [PyTorch installation guide](https://pytorch.org/get-started/locally/). They are required for this CPU example as well as for GPU inference; installing `funasr` alone does not select them for you. Then install the server dependencies in the same environment: + +```bash +pip install "funasr>=1.3.26" fastapi uvicorn python-multipart +``` + +Start the server on CPU: + +```bash +funasr-server \ + --host 0.0.0.0 \ + --port 8000 \ + --device cpu \ + --model sensevoice +``` + +Model weights must be cached locally or downloadable when the service loads the model. For NVIDIA GPU inference, select compatible CUDA builds in the PyTorch installation step above, then use `--device cuda`. + + +The command above binds all network interfaces so a LibreChat container can reach the host service. FunASR does not require authentication by default. Do not expose this endpoint directly to the public internet: restrict network access and use an authenticated TLS reverse proxy when remote access is required. If both processes run directly on the same host, use `--host 127.0.0.1` instead. + + +Verify the endpoint before configuring LibreChat: + +```bash +curl --fail http://localhost:8000/health + +curl --fail http://localhost:8000/v1/audio/transcriptions \ + -F file=@sample.wav \ + -F model=sensevoice \ + -F language=zh +``` + +A successful request returns JSON containing a `text` field. + +Then configure LibreChat: + +```yaml +speech: + stt: + allowedAddresses: + - 'host.docker.internal:8000' + openai: + url: 'http://host.docker.internal:8000/v1/audio/transcriptions' + model: 'sensevoice' +``` + +FunASR does not require an API key by default. Add `apiKey` only when an authenticated reverse proxy protects the endpoint. + +LibreChat blocks requests to private and loopback addresses unless the exact host and port are exempted under `speech.stt.allowedAddresses`. Use bare `host:port` entries without a scheme or path, and trust only hosts you control. See [Speech SSRF protection](/docs/configuration/librechat_yaml/object_structure/speech#ssrf-protection). + +Set both the URL and its matching exemption for your deployment; keep only the entry you need: + +Keep the URL scheme `http`, port `8000`, and path `/v1/audio/transcriptions` from the example above. Set the URL host and the single `allowedAddresses` entry as follows: + +- LibreChat in Docker and FunASR on the Docker host: URL host `host.docker.internal`, entry `host.docker.internal:8000` +- Both processes directly on the same host: URL host `127.0.0.1`, entry `127.0.0.1:8000` +- Both services in the same Compose network: for a service named `funasr`, URL host `funasr`, entry `funasr:8000`; replace the service name and port in both values if they differ + +On Linux, add a `host.docker.internal:host-gateway` mapping to the LibreChat container if that hostname is not already available. ## TTS (Text-to-Speech)