You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Open PR #732 adds a voice-input button to the composer. It only covers speech → prompt. In its current state it:
hardcodes engine: "gemini", which points at a local endpoint (127.0.0.1:8318) with a credential embedded in the source, so it won't work for other users;
has no settings UI for choosing an engine or entering a Whisper endpoint or key;
only ships en / zh-CN strings, with one status string hardcoded in Chinese;
includes unrelated changes: createUpdaterArtifacts: false in tauri.conf.json and a new build-windows.yml workflow.
Its UX ideas are good: a pulsing indicator while recording, a live preview of the in-progress transcript, and Esc to cancel. I'd reuse them. I'd rather not duplicate work, though, so tell me whether you want a fresh implementation that covers both directions or would rather see #732 reworked first.
Proposed behavior
Speech → prompt
A mic button in the composer. Click to start and click to stop, with a keyboard shortcut in Settings → Shortcuts.
A live transcript preview while speaking. The final text goes in at the caret and is never sent automatically. Esc cancels.
The recognition language defaults to the UI locale and can be overridden.
Agent response → speech
A "Read aloud" action on each finished agent message, with play/stop.
An optional "auto-read agent replies when a turn completes" toggle, off by default.
Speech covers the prose only. Code blocks, tool calls, diffs and raw markdown syntax are skipped or summarized ("code block omitted").
Playback stops when the user starts typing, sends a new prompt or switches sessions.
Settings (a new "Speech" section, or under General)
Engine selection for each direction. Everything is off by default, so nothing changes for existing users.
A configurable OpenAI-compatible endpoint, key and model for cloud engines (e.g. Whisper /v1/audio/transcriptions, /v1/audio/speech), stored the way codeg already stores provider credentials.
Voice, rate and language options.
Strings in all 10 supported locales.
Both runtimes
Desktop (Tauri) and server/web mode behave the same wherever the platform allows. Where an engine is unavailable, the UI says so and doesn't fail silently.
Engines: browser speech APIs (SpeechRecognition / speechSynthesis) are free and need no setup, but availability varies across Tauri's webviews. Speech recognition in particular isn't reliably available on WebKitGTK (Linux) or WebView2. Are you OK with:
(a) a browser engine where available, plus an OpenAI-compatible cloud engine configured by the user, or
(b) also adding an offline/native engine on the Rust side (e.g. local whisper), which increases binary size and build complexity?
Placement: a dedicated Settings → Speech page, or a section inside General?
PR shape: one PR, or two (speech → prompt first, then read-aloud)?
I'll start implementation once you've confirmed the direction.
Motivation
Codeg has no voice path today. Two directions would help, especially when you're away from the keyboard or reviewing long agent output:
I'd like to contribute both, and I'm opening this issue first to agree on scope with you before writing code.
Relation to #732
Open PR #732 adds a voice-input button to the composer. It only covers speech → prompt. In its current state it:
engine: "gemini", which points at a local endpoint (127.0.0.1:8318) with a credential embedded in the source, so it won't work for other users;en/zh-CNstrings, with one status string hardcoded in Chinese;createUpdaterArtifacts: falseintauri.conf.jsonand a newbuild-windows.ymlworkflow.Its UX ideas are good: a pulsing indicator while recording, a live preview of the in-progress transcript, and Esc to cancel. I'd reuse them. I'd rather not duplicate work, though, so tell me whether you want a fresh implementation that covers both directions or would rather see #732 reworked first.
Proposed behavior
Speech → prompt
Agent response → speech
Settings (a new "Speech" section, or under General)
/v1/audio/transcriptions,/v1/audio/speech), stored the way codeg already stores provider credentials.Both runtimes
Open questions for you
SpeechRecognition/speechSynthesis) are free and need no setup, but availability varies across Tauri's webviews. Speech recognition in particular isn't reliably available on WebKitGTK (Linux) or WebView2. Are you OK with:I'll start implementation once you've confirmed the direction.