Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 16 additions & 2 deletions .env.template
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,23 @@ AZURE_OPENAI_MODEL=
AZUREAI_API_KEY=
AZUREAI_REGION=
################################################################################
### ELEVENLABS API
### TTS PROVIDER
## Which text-to-speech provider to use: elevenlabs | sixtydb
## Keys below are only required for the selected provider.
################################################################################
TTS_PROVIDER=elevenlabs
################################################################################
### ELEVENLABS API (used when TTS_PROVIDER=elevenlabs)
## Eleven Labs Default Voice IDs
## https://elevenlabs.io/docs/voicelab/pre-made-voices
################################################################################
ELEVENLABS_API_KEY=
ELEVENLABS_VOICE_ID=
ELEVENLABS_VOICE_ID=
################################################################################
### 60DB API (used when TTS_PROVIDER=sixtydb)
## Dashboard: https://app.60db.ai | Docs: https://docs.60db.ai
## Voices: GET https://api.60db.ai/voices (Bearer auth)
## 60db default voice: fbb75ed2-975a-40c7-9e06-38e30524a9a1
################################################################################
SIXTYDB_API_KEY=
SIXTYDB_VOICE_ID=
35 changes: 28 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
# AI Assistant with fast Speech-to-Speech

This project is a Text or Voice Input AI Assistant that uses the OpenAI API to generate responses and the ElevenLabs Text-to-Speech API to convert the responses into audio quickly.
This project is a Text or Voice Input AI Assistant that uses the OpenAI API to generate responses and the ElevenLabs or 60db Text-to-Speech APIs to convert the responses into audio quickly.

## Features

- Real-time speech-to-speech conversation
- Customizable voice selection
- Switchable TTS provider: ElevenLabs or 60db (`TTS_PROVIDER` in `.env`)
- Easy-to-use command-line interface
- Powered by OpenAI and ElevenLabs
- **Paid ElevenLabs subscription required**
- Powered by OpenAI, ElevenLabs and 60db

# Getting Started

Expand All @@ -18,7 +18,9 @@ This project is a Text or Voice Input AI Assistant that uses the OpenAI API to g

- Get your [OpenAI API Key](https://platform.openai.com/api-keys)

- Get your [ElevenLabs API Key](https://elevenlabs.io/app/subscription)
- Get your [ElevenLabs API Key](https://elevenlabs.io/app/subscription) **or** your [60db API Key](https://app.60db.ai)

- The 60db TTS playback uses [mpv](https://mpv.io/installation/); install it and make sure it is on your PATH

- To run ```main_ws.py``` you will need a [Azure](https://portal.azure.com/) OpenAI Deployment and Speech Service Resource

Expand Down Expand Up @@ -50,7 +52,7 @@ pip install -r requirements.txt
cp .env.template .env
```

- Edit the `.env` file and add your OpenAI API key, ElevenLabs API key, and desired ElevenLabs voice ID.
- Edit the `.env` file and add your OpenAI API key plus the keys of the TTS provider you want to use (ElevenLabs or 60db).

## Usage

Expand All @@ -74,9 +76,27 @@ cp .env.template .env

## Customization

You can customize the voice used by the AI Assistant by changing the `ELEVENLABS_VOICE_ID` in the `.env` file.
### Choosing the TTS provider

Set `TTS_PROVIDER` in the `.env` file to switch between providers:

```bash
TTS_PROVIDER=elevenlabs # ElevenLabs (default)
TTS_PROVIDER=sixtydb # 60db
```

Only the API keys of the selected provider are required:

- `elevenlabs`: `ELEVENLABS_API_KEY`, `ELEVENLABS_VOICE_ID`
- `sixtydb`: `SIXTYDB_API_KEY`, `SIXTYDB_VOICE_ID`

`main.py` uses the 60db TTS stream API (`POST /tts-stream`), `main_ws.py` uses the 60db TTS WebSocket (`wss://api.60db.ai/ws/tts`).

### Changing the voice

You can customize the voice used by the AI Assistant by changing the `ELEVENLABS_VOICE_ID` or `SIXTYDB_VOICE_ID` in the `.env` file.

A link to the pre-made voices is in the `.env.template` file.
Links to the pre-made voices are in the `.env.template` file.

## License

Expand All @@ -97,3 +117,4 @@ If you encounter any issues or have questions, please open an issue on GitHub.

- [OpenAI](https://www.openai.com/)
- [ElevenLabs](https://www.elevenlabs.io/)
- [60db](https://60db.ai/)
23 changes: 18 additions & 5 deletions config.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,10 +32,23 @@
if OPENAI_SYSTEM_PROMPT is None:
raise ValueError("OPENAI_SYSTEM_PROMPT not set")

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
if ELEVENLABS_API_KEY is None:
raise ValueError("ELEVENLABS_API_KEY not set")
TTS_PROVIDER = os.getenv("TTS_PROVIDER", "elevenlabs").strip().lower()
if TTS_PROVIDER not in ("elevenlabs", "sixtydb"):
raise ValueError("TTS_PROVIDER must be 'elevenlabs' or 'sixtydb'")

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
ELEVENLABS_VOICE_ID = os.getenv("ELEVENLABS_VOICE_ID")
if ELEVENLABS_VOICE_ID is None:
raise ValueError("ELEVENLABS_VOICE_ID not set")

SIXTYDB_API_KEY = os.getenv("SIXTYDB_API_KEY")
SIXTYDB_VOICE_ID = os.getenv("SIXTYDB_VOICE_ID")

if TTS_PROVIDER == "elevenlabs":
if ELEVENLABS_API_KEY is None:
raise ValueError("ELEVENLABS_API_KEY not set")
if ELEVENLABS_VOICE_ID is None:
raise ValueError("ELEVENLABS_VOICE_ID not set")
else:
if SIXTYDB_API_KEY is None:
raise ValueError("SIXTYDB_API_KEY not set")
if SIXTYDB_VOICE_ID is None:
raise ValueError("SIXTYDB_VOICE_ID not set")
23 changes: 18 additions & 5 deletions config_ws.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,10 +40,23 @@
if AZUREAI_REGION is None:
raise ValueError("AZUREAI_REGION not set")

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
if ELEVENLABS_API_KEY is None:
raise ValueError("ELEVENLABS_API_KEY not set")
TTS_PROVIDER = os.getenv("TTS_PROVIDER", "elevenlabs").strip().lower()
if TTS_PROVIDER not in ("elevenlabs", "sixtydb"):
raise ValueError("TTS_PROVIDER must be 'elevenlabs' or 'sixtydb'")

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
ELEVENLABS_VOICE_ID = os.getenv("ELEVENLABS_VOICE_ID")
if ELEVENLABS_VOICE_ID is None:
raise ValueError("ELEVENLABS_VOICE_ID not set")

SIXTYDB_API_KEY = os.getenv("SIXTYDB_API_KEY")
SIXTYDB_VOICE_ID = os.getenv("SIXTYDB_VOICE_ID")

if TTS_PROVIDER == "elevenlabs":
if ELEVENLABS_API_KEY is None:
raise ValueError("ELEVENLABS_API_KEY not set")
if ELEVENLABS_VOICE_ID is None:
raise ValueError("ELEVENLABS_VOICE_ID not set")
else:
if SIXTYDB_API_KEY is None:
raise ValueError("SIXTYDB_API_KEY not set")
if SIXTYDB_VOICE_ID is None:
raise ValueError("SIXTYDB_VOICE_ID not set")
95 changes: 91 additions & 4 deletions main.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,14 @@
"""
This script uses the OpenAI API and ElevenLab's REST API
This script uses the OpenAI API and the ElevenLabs REST API
or the 60db TTS stream API, selected via TTS_PROVIDER.
"""
import base64
import json
import shutil
import subprocess
import sys

import requests
import speech_recognition as sr # type: ignore
from openai import OpenAI
from elevenlabs import stream
Expand All @@ -14,8 +21,11 @@
OPENAI_PROJECT_ID,
OPENAI_MODEL,
OPENAI_SYSTEM_PROMPT,
TTS_PROVIDER,
ELEVENLABS_API_KEY,
ELEVENLABS_VOICE_ID,
SIXTYDB_API_KEY,
SIXTYDB_VOICE_ID,
)

console = Console()
Expand All @@ -25,14 +35,15 @@
organization=OPENAI_ORG_ID,
project=OPENAI_PROJECT_ID,
)
e_client = ElevenLabs(api_key=ELEVENLABS_API_KEY)
if TTS_PROVIDER == "elevenlabs":
e_client = ElevenLabs(api_key=ELEVENLABS_API_KEY)
voice_ident = ELEVENLABS_VOICE_ID


def generate_and_play_response(user_input, conversation_history):
"""
Generates response using the OpenAI API and
plays it using the ElevenLabs TTS API.
Generates response using the OpenAI API and plays it using
the selected TTS provider (ElevenLabs or 60db).

Args:
user_input (str): The user's input.
Expand Down Expand Up @@ -72,6 +83,10 @@ def generate_and_play_response(user_input, conversation_history):
def text_stream():
yield response_text

if TTS_PROVIDER == "sixtydb":
tts_stream_sixtydb(response_text)
return

audio_stream = e_client.generate(
text=text_stream(),
voice=Voice(
Expand All @@ -91,6 +106,78 @@ def text_stream():
stream(audio_stream)


def tts_stream_sixtydb(text):
"""
Converts text to speech using the 60db TTS stream API
(POST https://api.60db.ai/tts-stream) and plays the streamed
MP3 chunks through the mpv player.

The response is newline-delimited JSON (NDJSON): each "chunk"
message contains a base64-encoded piece of audio, followed by a
final "complete" message.

Args:
text (str): The text to convert to speech.

Returns:
None

Raises:
ValueError: If mpv is not installed on the system.
RuntimeError: If the 60db API returns an error message.
"""
if shutil.which("mpv") is None:
raise ValueError(
"mpv not found, necessary to stream audio. "
"Install instructions: https://mpv.io/installation/"
)

response = requests.post(
"https://api.60db.ai/tts-stream",
headers={
"Authorization": f"Bearer {SIXTYDB_API_KEY}",
"Content-Type": "application/json",
},
json={
"text": text,
"voice_id": SIXTYDB_VOICE_ID,
"speed": 1,
"stability": 50,
"similarity": 75,
},
stream=True,
timeout=60,
)
response.raise_for_status()

mpv_process = subprocess.Popen(
["mpv", "--no-cache", "--no-terminal", "--", "fd://0"],
stdin=subprocess.PIPE,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
)

try:
for line in response.iter_lines():
if not line:
continue
data = json.loads(line)
if data.get("type") == "chunk":
audio_chunk = base64.b64decode(
data["result"]["audioContent"]
)
mpv_process.stdin.write(audio_chunk)
mpv_process.stdin.flush()
elif data.get("type") == "error":
raise RuntimeError(
f"60db TTS error: {data.get('message')}"
)
finally:
if mpv_process.stdin:
mpv_process.stdin.close()
mpv_process.wait()


def recognize_speech(timeout=20):
"""
Recognizes speech from the microphone input.
Expand Down
Loading