Skip to content

feat: whisper.cpp 1.9 API catch-up - #15

Draft
rmorse wants to merge 6 commits into
developfrom
feat/whisper-api-catchup
Draft

rmorse wants to merge 6 commits into
developfrom
feat/whisper-api-catchup

Conversation

@rmorse

@rmorse rmorse commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Draft PR for Phase 1 of the whisper.cpp 1.9 upgrade: exposing whisper.cpp APIs the safe crate has missed over recent release cycles, plus the foundations they depend on. Commits are added as each item lands; it will be marked ready and squash-merged when Phase 1 is complete.

Release notes (draft)

User-facing summary of everything in this PR, kept current so it can be reused for the release notes. Mirrors the CHANGELOG.md [Unreleased] entries added here.

Added

  • WhisperLog: control whisper.cpp's log output (model loading, processing, VAD and ggml backend messages), wrapping whisper_log_set.
    • WhisperLog::set(|level, message| ...) routes messages to your own callback with a LogLevel (Debug, Info, Warn, Error).
    • WhisperLog::disable() silences whisper.cpp entirely.
    • WhisperLog::reset() restores the default stderr output.
    • All three are safe to call at any time, including during transcription. A log hook set directly through whisper-cpp-plus-sys before the crate's first call into whisper.cpp is replaced at that call.
    • Output still goes to stderr by default; nothing changes unless you opt in.
  • log feature: WhisperLog::use_log_crate() forwards whisper.cpp output to the log crate (target whisper_cpp), so env_logger, tracing-log and similar work out of the box.
  • Streaming Silero VAD: WhisperVadProcessor::detect_speech_no_reset() and reset_state() (wrapping whisper_vad_detect_speech_no_reset / whisper_vad_reset_state) evaluate consecutive chunks of a stream in context, plus WhisperVadProcessor::WINDOW_SAMPLES (512 samples per probability).
  • Parallel transcription that works: WhisperContext::full_parallel(params, audio, n_processors) splits audio into equal chunks, transcribes them concurrently, and returns the merged TranscriptionResult with times on the original timeline. Chunking follows whisper.cpp's whisper_full_parallel, with one improvement: segment times are clamped to their chunk, so an overrunning segment end can't push the next chunk's segments later. Words that straddle a chunk boundary may still be cut or misrecognised.
  • WhisperState::n_len(): mel length of the last transcription on a state.

Changed

  • Lower memory per model: WhisperContext no longer allocates whisper.cpp's default state, which the crate never used. This saves about 146 MB per loaded context with ggml-tiny.en.bin (as reported by whisper.cpp), and considerably more for larger models.
  • The whisper-cpp-plus-sys package now includes whisper.cpp's public headers and license, so docs.rs generates bindings from the real headers. Regular builds are unchanged (they still download the full pinned whisper.cpp source).

Removed (breaking)

  • WhisperState::full_parallel(): it never returned correct results (whisper.cpp wrote them to the context's default state, which the method never read). Use WhisperContext::full_parallel().
  • WhisperContext::n_len(): it reported the context's unused default state. Use WhisperState::n_len().

Fixed

  • More accurate Silero VAD in WhisperStreamPcm. Each 200 ms probe used to be judged by a freshly reset model, so speech onsets and short words were often misclassified: on jfk.wav, 16 of 55 probe decisions differed from a full-file Silero pass, cutting "Ask not" short (transcribed as "Ask, knock!") and splitting a sentence in two. The model state is now carried across the whole stream, which matches the full-file pass; the same clip now transcribes as "Ask not!" with segments that follow the speaker's pauses.
  • whisper-cpp-plus-sys documentation on docs.rs now matches the real API. It was generated from out-of-date hand-written stubs with missing functions, nonexistent functions and some wrong signatures.

Documentation

  • WhisperVadProcessor::detect_speech() returns whether the computation succeeded, not whether speech was found; probabilities come from get_probs().
  • PcmReaderConfig::buffer_len_ms drops the oldest samples when full. Sources faster than real time (files, in-memory buffers) need a buffer that holds the whole input.

Implementation notes

  • docs.rs bindings (f584ad8): hand-written stubs in build.rs (~300 lines) replaced by the normal bindgen step against packaged headers (include/*.h, ggml/include/*.h, 34 files / 66 KiB compressed). The docs.rs image provides libclang. The bundled-source check now requires CMakeLists.txt + src/whisper.cpp, so the headers-only package still downloads the pinned commit. New CI step simulates the docs.rs build from the packaged .crate.

  • Logging (43379f3, 88c1645): whisper.cpp only holds a pointer to one trampoline. whisper_log_set writes whisper.cpp's global log state unsynchronised while logging reads it, so the trampoline is installed exactly once (std::sync::Once) before the crate's first call into whisper.cpp (context/VAD constructors and quantization call logging::ensure_installed()) and never changed again; set/disable/reset only swap the Rust-side sink in a static RwLock. (43379f3 installed on the first WhisperLog call and uninstalled on reset, which raced with in-flight transcriptions; found by the external design review.) catch_unwind in the trampoline (MSRV 1.70: unwinding into C is UB). Levels compared via bindgen constants (ggml_log_level is c_int on MSVC, c_uint elsewhere). GGML_LOG_LEVEL_CONT mapped to the previous level (only emitted by Vulkan/OpenCL). Integration tests run in their own binaries because the hook is process-global.

  • Streaming VAD (a303ebd): WhisperStreamPcm resets the VAD once when the stream is created, feeds only whole 512-sample windows via detect_speech_no_reset and carries the remainder; a probe that doesn't complete a window keeps the previous decision. Deliberately does not reset between utterances (the upstream header suggests it) because that restarts the model cold at each speech onset. New test asserts streamed probabilities equal a full-file pass (< 1e-4) and prints the old method's error for comparison.

  • Stream test fix (a303ebd): the stream_pcm_integration tests read the 11 s jfk.wav into a 10 s PcmReader buffer faster than real time, so the reader dropped the first second and the tests never transcribed the opening words. Buffers are now sized from the clip, and the keyword check asserts "and so" is present.

  • Parallel transcription + no default state (eca1efa): full_parallel uses std::thread::scope with one WhisperState per chunk on a shared context (as upstream does); the first chunk keeps offset_ms, later chunks start after it. Chunk results are clamped to chunk bounds before the no-overlap merge (upstream doesn't clamp; with jfk.wav split in two, an unclamped chunk end overran by ~1.4 s). With full_parallel and n_len no longer touching the default state, contexts load via whisper_init_*_with_params_no_state; a test asserts (via WhisperLog) that loading a context triggers no whisper_init_state allocation and creating a WhisperState does.

  • Pin description corrected (0263fd5): the pinned commit contains the upstream v1.9.4 release plus 181 later commits. The READMEs, CHANGELOG and build comments (merged in chore: update whisper.cpp pin to 1.9.4-dev (stream-pcm de8fb5fd) #13) described it as "1.9.4-dev, based on master after v1.9.3"; they now say post-v1.9.4, and the CHANGELOG's release range is v1.8.7 through v1.9.4.

Planned

  • API layers redesign (under review): mirror both whisper.cpp API layers (default-state context layer and multi-state layer) with upstream's defaults. This will revise eca1efa: new allocates the default state again (as 0.1.5 did) with new *_no_state constructors, WhisperContext::n_len returns, full_parallel becomes the upstream mirror, and the Rust implementation is renamed transcribe_parallel.
  • Context params: use_gpu, gpu_device, flash_attn, DTW token timestamps
  • Callbacks: abort/cancellation, new segment, progress, encoder begin
  • Language detection and tokenization helpers
  • Remaining FullParams setters and model info

Validation (so far)

Windows / MSVC, test models from cargo xtask test-setup:

  • cargo fmt --all -- --check; clippy (default, async, log) with -D warnings
  • cargo test --workspace -- --test-threads=1: 135 passed, none skipped; async suite: 123 passed
  • --features log: logging_log_crate passes; cargo doc --features log clean
  • docs.rs simulation from the extracted .crate (headers only, offline): 134 bindgen functions; safe crate documents cleanly with --all-features
  • cargo package -p whisper-cpp-plus-sys verify build: downloads the pinned de8fb5fd source and builds (crates.io consumer path)
  • Streaming VAD before/after on jfk.wav (corrected test buffers): old 4 segments incl. "Ask, knock!"; new 3 segments, "Ask not!"
  • full_parallel on jfk.wav: 2 chunks merge in order on the original timeline, clamped at the split; offset_ms respected; n_processors = 1 and too-short audio match a single transcription

macOS CI: green on f584ad8 (including "Check docs.rs build" and Metal tests), 43379f3 and a303ebd.

Replace the hand-written docs.rs stub bindings, which had drifted from whisper.h (missing and nonexistent functions, wrong signatures and types), with the normal bindgen step run against whisper.cpp's public headers. The sys crate package now ships include/*.h, ggml/include/*.h and whisper.cpp's LICENSE; the docs.rs image provides libclang.

The bundled-source check now requires the full source tree, so builds of the headers-only package still download the pinned commit. CI simulates the docs.rs build from the packaged crate.
Wrap whisper_log_set, which also routes ggml backend and VAD logs:
WhisperLog::set() sends messages to a Rust callback with a LogLevel,
disable() silences them and reset() restores whisper.cpp's stderr default.
The optional `log` feature adds WhisperLog::use_log_crate(), forwarding to
the log crate with target "whisper_cpp".

whisper.cpp only ever holds a pointer to a single trampoline, installed
once; the Rust callback lives in a static and can be swapped at any time.
Callback panics are caught so they never unwind into C.
WhisperStreamPcm evaluated every 200 ms probe with whisper_vad_detect_speech,
which resets Silero's recurrent state, and the zero-padded partial 512-sample
window at the end of each probe skewed the result. On jfk.wav, 16 of 55 probe
decisions differed from a full-file Silero pass, cutting "Ask not" short
("Ask, knock!") and splitting a sentence.

Wrap whisper_vad_detect_speech_no_reset / whisper_vad_reset_state as
WhisperVadProcessor::detect_speech_no_reset() / reset_state() and add
WINDOW_SAMPLES. WhisperStreamPcm now resets once per stream, feeds only whole
windows and carries the remainder, matching the full-file pass. A new test
asserts streamed probabilities equal the full pass.

Also fix the stream_pcm integration tests: their 10 s PcmReader buffer was
shorter than the 11 s clip, so the reader (correctly, for live input) dropped
the first second. Buffers are now sized from the clip and the tests check the
opening words; PcmReaderConfig::buffer_len_ms documents the overflow policy.
…state

WhisperState::full_parallel called whisper_full_parallel, which writes its
results to the context's default state; the method then read its own state,
so it never returned correct segments. Replace it with
WhisperContext::full_parallel(params, audio, n_processors) -> TranscriptionResult,
which follows whisper.cpp's chunking (equal chunks after offset_ms, one state
per chunk, times shifted onto the original timeline, no overlap) using scoped
threads. Segment times are also clamped to their chunk, so a segment end that
whisper reports past its audio can't push the next chunk's segments later.

With no remaining default-state users, WhisperContext now loads with
whisper_init_*_with_params_no_state, saving a full state's KV caches and
compute buffers per context (~146 MB for tiny.en). WhisperContext::n_len,
which read that default state, moves to WhisperState::n_len.

BREAKING CHANGE: WhisperState::full_parallel and WhisperContext::n_len are
removed; use WhisperContext::full_parallel and WhisperState::n_len.
The pinned commit contains the upstream v1.9.4 release plus 181 later upstream commits; the READMEs, CHANGELOG, Cargo.toml and build.rs comments described it as 1.9.4-dev based on master after v1.9.3.
whisper_log_set writes whisper.cpp's global log state without
synchronisation while logging reads it, so calling it from WhisperLog
while another thread was inside whisper.cpp was a data race. The crate
now installs its trampoline exactly once (std::sync::Once) before its
first call into whisper.cpp: context and VAD constructors and the
quantization entry points call logging::ensure_installed(). WhisperLog
set/disable/reset only change the Rust-side sink afterwards, so they are
safe at any time. reset() now writes messages to stderr unchanged, as
happens when no hook has been set.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant