Skip to content

[Feature] Speech support: voice input to prompt and read-aloud of agent responses #844

Description

@n0tlu5

Motivation

Codeg has no voice path today. Two directions would help, especially when you're away from the keyboard or reviewing long agent output:

  1. Speech → prompt: dictate into the chat composer instead of typing.
  2. Agent response → speech: have an agent's reply read aloud.

I'd like to contribute both, and I'm opening this issue first to agree on scope with you before writing code.

Relation to #732

Open PR #732 adds a voice-input button to the composer. It only covers speech → prompt. In its current state it:

  • hardcodes engine: "gemini", which points at a local endpoint (127.0.0.1:8318) with a credential embedded in the source, so it won't work for other users;
  • has no settings UI for choosing an engine or entering a Whisper endpoint or key;
  • only ships en / zh-CN strings, with one status string hardcoded in Chinese;
  • includes unrelated changes: createUpdaterArtifacts: false in tauri.conf.json and a new build-windows.yml workflow.

Its UX ideas are good: a pulsing indicator while recording, a live preview of the in-progress transcript, and Esc to cancel. I'd reuse them. I'd rather not duplicate work, though, so tell me whether you want a fresh implementation that covers both directions or would rather see #732 reworked first.

Proposed behavior

Speech → prompt

  • A mic button in the composer. Click to start and click to stop, with a keyboard shortcut in Settings → Shortcuts.
  • A live transcript preview while speaking. The final text goes in at the caret and is never sent automatically. Esc cancels.
  • The recognition language defaults to the UI locale and can be overridden.

Agent response → speech

Settings (a new "Speech" section, or under General)

  • Engine selection for each direction. Everything is off by default, so nothing changes for existing users.
  • A configurable OpenAI-compatible endpoint, key and model for cloud engines (e.g. Whisper /v1/audio/transcriptions, /v1/audio/speech), stored the way codeg already stores provider credentials.
  • Voice, rate and language options.
  • Strings in all 10 supported locales.

Both runtimes

  • Desktop (Tauri) and server/web mode behave the same wherever the platform allows. Where an engine is unavailable, the UI says so and doesn't fail silently.

Open questions for you

  1. feat(chat): add voice input to message composer with dual-mode recognition #732: fresh implementation, or build on feat(chat): add voice input to message composer with dual-mode recognition #732?
  2. Engines: browser speech APIs (SpeechRecognition / speechSynthesis) are free and need no setup, but availability varies across Tauri's webviews. Speech recognition in particular isn't reliably available on WebKitGTK (Linux) or WebView2. Are you OK with:
    • (a) a browser engine where available, plus an OpenAI-compatible cloud engine configured by the user, or
    • (b) also adding an offline/native engine on the Rust side (e.g. local whisper), which increases binary size and build complexity?
  3. Placement: a dedicated Settings → Speech page, or a section inside General?
  4. PR shape: one PR, or two (speech → prompt first, then read-aloud)?

I'll start implementation once you've confirmed the direction.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions