Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .github/workflows/macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -74,11 +74,29 @@ jobs:
- name: Check async Clippy
run: cargo clippy -p whisper-cpp-plus --all-targets --features async -- -D warnings

- name: Check log Clippy
run: cargo clippy -p whisper-cpp-plus --all-targets --features log -- -D warnings

- name: Check Metal Clippy
run: cargo clippy -p whisper-cpp-plus --all-targets --features metal -- -D warnings
env:
MACOSX_DEPLOYMENT_TARGET: "14.0"

# docs.rs has no network access and builds from the published package, which only ships
# whisper.cpp's public headers. Simulate that so packaging/header regressions fail here.
- name: Check docs.rs build
env:
DOCS_RS: "1"
run: |
cargo package -p whisper-cpp-plus-sys --no-verify
mkdir -p "$RUNNER_TEMP/docsrs"
tar -xzf target/package/whisper-cpp-plus-sys-*.crate -C "$RUNNER_TEMP/docsrs"
cargo doc --no-deps \
--manifest-path "$(echo "$RUNNER_TEMP"/docsrs/whisper-cpp-plus-sys-*/Cargo.toml)" \
--target-dir "$RUNNER_TEMP/docsrs/target"
cargo doc --no-deps -p whisper-cpp-plus --all-features \
--target-dir "$RUNNER_TEMP/docsrs/target"

- name: Cache test models
uses: actions/cache@v5
with:
Expand Down Expand Up @@ -133,6 +151,9 @@ jobs:
- name: Test workspace
run: cargo test --workspace -- --test-threads=1

- name: Test log feature
run: cargo test -p whisper-cpp-plus --features log --test logging_log_crate -- --test-threads=1

- name: Test Metal feature
run: cargo test -p whisper-cpp-plus --features metal -- --test-threads=1
env:
Expand Down
21 changes: 20 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,23 +9,42 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Changed

- Updated the pinned whisper.cpp fork to `rmorse/whisper.cpp` `stream-pcm` at `de8fb5fd` (tag `v1.9.4-dev-stream-pcm`, whisper.cpp `1.9.4-dev`), based on upstream `ggml-org/whisper.cpp` `master` after `v1.9.3`. This picks up upstream releases `v1.8.7` through `v1.9.3` and ggml `0.25.1`.
- Updated the pinned whisper.cpp fork to `rmorse/whisper.cpp` `stream-pcm` at `de8fb5fd` (tag `v1.9.4-dev-stream-pcm`), based on upstream `ggml-org/whisper.cpp` `master` after the `v1.9.4` release (`v1.9.4` plus 181 later upstream commits). This picks up upstream releases `v1.8.7` through `v1.9.4` and ggml `0.25.1`.
- Upstream now re-seeds the decoder between calls (ggml-org/whisper.cpp#4025), so temperature-fallback output is deterministic across repeated transcriptions on the same state.
- Upstream now rejects Silero VAD models whose encoder does not have exactly 4 layers (ggml-org/whisper.cpp#4064); loading such a model with `WhisperVadProcessor` now fails at load time.
- On Apple Silicon, upstream's optional ANEForge encoder backend (ggml-org/whisper.cpp#3905) is activated by the `ANEFORGE_ENCODER` and `ANEFORGE_DYLIB` environment variables, which load a dynamic library from the given path when a state is created. It is inactive unless those variables are set.
- NVIDIA Parakeet support is available in the bundled C library but is not yet exposed through the Rust API.
- The new upstream VAD segment and VAD-mapped token timestamp accessors are not exposed: they are only populated by upstream's built-in `whisper_full` VAD, which does not run for per-state transcription (`whisper_full_with_state`, used by this crate; see ggml-org/whisper.cpp#3423). Use the crate's own VAD pipeline instead.
- The `whisper-cpp-plus-sys` package now includes whisper.cpp's public headers (`include/*.h`, `ggml/include/*.h`) and its `LICENSE`. docs.rs builds generate bindings from these headers instead of using hand-written stubs. Regular builds are unchanged: they still download the full pinned whisper.cpp source.
- `WhisperContext` no longer allocates whisper.cpp's default state, which the crate never used (all transcription runs on explicit `WhisperState`s). This saves that state's KV caches and compute buffers for every loaded context: about 146 MB with `ggml-tiny.en.bin` as reported by whisper.cpp, and considerably more for larger models.

### Removed

- **Breaking:** `WhisperState::full_parallel()`. It never returned correct results: whisper.cpp writes parallel results to the context's default state, which the method never read. Use `WhisperContext::full_parallel()`, which returns the merged `TranscriptionResult`.
- **Breaking:** `WhisperContext::n_len()`. It reported the mel length of the context's default state, which the crate never transcribes on. Use `WhisperState::n_len()` for the state you transcribed with.

### Added

- `WhisperState::full_get_segment_no_speech_prob()`, previously only used internally by the temperature-fallback transcriber.
- `WhisperLog` (wrapping `whisper_log_set`) to control whisper.cpp's log output, which covers whisper.cpp, its VAD and the ggml backends: `WhisperLog::set()` routes messages to a Rust callback with a `LogLevel`, `WhisperLog::disable()` silences them, and `WhisperLog::reset()` restores the default stderr output. All three are safe to call at any time, including during transcription: whisper.cpp's log hook is unsynchronised global state, so the crate installs its own hook once, before its first call into whisper.cpp, and `WhisperLog` only changes where that hook sends messages. A hook set directly through `whisper-cpp-plus-sys` before the crate's first call is replaced.
- `log` feature: `WhisperLog::use_log_crate()` forwards whisper.cpp log output to the `log` crate with target `whisper_cpp`.
- `WhisperVadProcessor::detect_speech_no_reset()` and `reset_state()` (wrapping `whisper_vad_detect_speech_no_reset` / `whisper_vad_reset_state`) for streaming Silero VAD that keeps its state across chunks, plus `WhisperVadProcessor::WINDOW_SAMPLES` (512 samples per probability).
- `WhisperContext::full_parallel(params, audio, n_processors)`: splits audio into equal chunks, transcribes them concurrently on separate states, and returns the merged `TranscriptionResult` with times on the original timeline. Chunking follows whisper.cpp's `whisper_full_parallel`; in addition, segment times are clamped to their chunk, so whisper reporting a segment end past its audio can no longer push the next chunk's segments later. Words that straddle a chunk boundary may still be cut or misrecognised.
- `WhisperState::n_len()`: mel length of the last transcription on the state (`whisper_n_len_from_state`).

### Fixed

- **Breaking:** segment timestamps are now real milliseconds. whisper.cpp reports segment times in centiseconds, and the crate previously passed them through unconverted, so `Segment::start_ms`/`end_ms`, `start_seconds()`/`end_seconds()`, `WhisperState::full_get_segment_timestamps()`, and the `start`/`end` values passed to `WhisperStreamPcm::run` callbacks were 10x too small. This affects `transcribe*`, `WhisperStream`, `WhisperStreamPcm`, and the temperature-fallback transcriber. The raw `whisper_token_data` returned by `full_get_token_data()` is unchanged and documented as centiseconds.
- Fixed a use-after-free in `FullParams::suppress_regex()`: the regex string was freed immediately after being set, so whisper.cpp read freed memory during transcription.
- Fixed `FullParams::prompt_tokens()` storing a borrowed pointer that dangled once the caller's slice was dropped or the params were moved or cloned. The tokens are now copied into the params.
- **Breaking:** `WhisperState` result getters now validate segment and token indices instead of passing them to whisper.cpp, which does not bounds-check (out-of-range indices were undefined behaviour). `full_get_segment_text()` and `full_get_token_text()` return `WhisperError::InvalidParameter`, `full_get_token_data()` returns `None`, and the plain-value getters (`full_get_segment_timestamps()`, `full_get_segment_speaker_turn_next()`, `full_n_tokens()`, `full_get_token_id()`, `full_get_token_prob()`) panic, like slice indexing.
- Improved Silero VAD accuracy in `WhisperStreamPcm`. Each 200 ms probe was evaluated from a freshly reset model with a zero-padded partial window, so speech onsets and short words were often misclassified: on `jfk.wav`, 16 of 55 probe decisions differed from a full-file Silero pass, cutting "Ask not" short (transcribed as "Ask, knock!") and splitting a sentence. The model state is now carried across probes (reset only when the stream starts) and only whole 32 ms windows are evaluated, which matches the full-file pass.
- Fixed the `whisper-cpp-plus-sys` documentation on docs.rs, which was generated from out-of-date hand-written stubs: it was missing functions, listed functions that no longer exist, and showed some wrong signatures and types. It now matches the real bindings.

### Documentation

- Clarified that `WhisperVadProcessor::detect_speech()` returns whether the computation succeeded, not whether speech was found; speech probabilities come from `get_probs()`.
- Documented that `PcmReaderConfig::buffer_len_ms` drops the oldest samples on overflow, so sources faster than real time (files, in-memory buffers) need a buffer that holds the whole input.

## [0.1.5] - 2026-06-12

Expand Down
7 changes: 7 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,13 @@ cargo clippy -p whisper-cpp-plus --all-targets --features async -- -D warnings
cargo test -p whisper-cpp-plus --features async -- --test-threads=1
```

For logging (`log` feature) changes:

```bash
cargo clippy -p whisper-cpp-plus --all-targets --features log -- -D warnings
cargo test -p whisper-cpp-plus --features log --test logging_log_crate -- --test-threads=1
```

For macOS or Metal-sensitive changes:

```bash
Expand Down
2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Pinned to whisper.cpp 1.9.4-dev (fork: rmorse/whisper.cpp, branch: stream-pcm, commit de8fb5fd)
# Pinned to whisper.cpp post-v1.9.4 (fork: rmorse/whisper.cpp, branch: stream-pcm, commit de8fb5fd)
[workspace]
members = ["whisper-cpp-plus-sys", "whisper-cpp-plus", "xtask"]
resolver = "2"
Expand Down
30 changes: 28 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# whisper-cpp-plus

> **Pinned to whisper.cpp 1.9.4-dev** (fork: [`rmorse/whisper.cpp`](https://github.com/rmorse/whisper.cpp), branch: `stream-pcm`, commit [`de8fb5fd`](https://github.com/rmorse/whisper.cpp/commit/de8fb5fda8b25837a2ba0034c8c24223a6fd6c6c), based on upstream `ggml-org/whisper.cpp` `master` after `v1.9.3`)
> **Pinned to whisper.cpp post-v1.9.4** (fork: [`rmorse/whisper.cpp`](https://github.com/rmorse/whisper.cpp), branch: `stream-pcm`, commit [`de8fb5fd`](https://github.com/rmorse/whisper.cpp/commit/de8fb5fda8b25837a2ba0034c8c24223a6fd6c6c), based on upstream `ggml-org/whisper.cpp` `master` after the `v1.9.4` release)

Safe Rust bindings for [whisper.cpp](https://github.com/ggerganov/whisper.cpp) with real-time PCM streaming and VAD support — OpenAI's Whisper speech recognition model.

Expand Down Expand Up @@ -35,6 +35,7 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
- **Async** — `tokio::spawn_blocking` wrappers (feature = `async`)
- **Cross-platform** — Windows (MSVC), Linux, macOS (Intel & Apple Silicon)
- **Quantization** — model compression via `WhisperQuantize` (feature = `quantization`)
- **Logging control** — redirect, silence, or forward whisper.cpp's log output via `WhisperLog` (`log` crate integration with feature = `log`)
- **Hardware acceleration** — SIMD auto-detected, GPU via feature flags

## Installation
Expand All @@ -60,6 +61,7 @@ whisper-cpp-plus = { version = "0.1.5", features = ["quantization"] } # Model q
whisper-cpp-plus = { version = "0.1.5", features = ["async"] } # Async API
whisper-cpp-plus = { version = "0.1.5", features = ["cuda"] } # NVIDIA GPU
whisper-cpp-plus = { version = "0.1.5", features = ["metal"] } # macOS GPU
whisper-cpp-plus = { version = "0.1.5", features = ["log"] } # Forward logs to the `log` crate
```

### CUDA GPU Acceleration
Expand Down Expand Up @@ -205,7 +207,7 @@ Notes:

- `PcmReader` does not decode WAV/MP3, resample audio, or convert stereo to mono. Your `Read` source must already be normalized to the format described by `PcmReaderConfig`.
- `WhisperStreamPcm::new(...)` uses fixed-step mode or simple built-in VAD depending on `use_vad`.
- `WhisperStreamPcm::with_vad(...)` uses an explicit `WhisperVadProcessor` (Silero VAD) and is the recommended path when you want Silero-based segmentation.
- `WhisperStreamPcm::with_vad(...)` uses an explicit `WhisperVadProcessor` (Silero VAD) and is the recommended path when you want Silero-based segmentation. Silero's state is carried across probes for the whole stream (it is reset when the stream is created), so each probe is judged in context.
- In VAD mode, `no_context` is forced internally to match `stream-pcm.cpp`.
- In VAD mode, `run()` emits the next completed speech chunk in chronological order, and callers can usually append those segments directly.
- In fixed-step mode, callbacks are produced from overlapping windows, so callers that build a cumulative transcript may need to reconcile repeated text across callbacks.
Expand Down Expand Up @@ -254,6 +256,30 @@ let result = ctx.transcribe_with_params_enhanced(&audio, params)?;
// Automatically retries with higher temperatures if quality thresholds aren't met
```

**Controlling whisper.cpp log output:**

whisper.cpp prints model-loading and processing details to stderr by default. `WhisperLog` changes where they go, and can be called at any time (usually once at startup):

```rust
use whisper_cpp_plus::{LogLevel, WhisperLog};

// Silence whisper.cpp entirely
WhisperLog::disable();

// ...or route messages to your own handler
WhisperLog::set(|level, message| {
if level >= LogLevel::Warn {
eprintln!("[whisper.cpp {:?}] {}", level, message);
}
});

// ...or, with feature = "log", forward to the `log` crate (target "whisper_cpp")
// WhisperLog::use_log_crate();

// Restore the default stderr output
WhisperLog::reset();
```

More examples in [`whisper-cpp-plus/examples/`](./whisper-cpp-plus/examples/).

## Enhanced Features
Expand Down
2 changes: 1 addition & 1 deletion docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,4 +65,4 @@ Zero-copy where possible: `&[f32]` slices passed directly to C++ via `.as_ptr()`

## Build system

`cmake` crate compiles whisper.cpp sources. whisper.cpp vendored as git submodule at `whisper-cpp-plus-sys/whisper.cpp` (excluded from crates.io package - downloaded on demand for consumers). Prebuilt library caching available via `cargo xtask prebuild`.
`cmake` crate compiles whisper.cpp sources. whisper.cpp vendored as git submodule at `whisper-cpp-plus-sys/whisper.cpp`. The crates.io package ships only its public headers and license; consumers download the full pinned source on demand, and docs.rs (no network) generates bindings from the packaged headers. Prebuilt library caching available via `cargo xtask prebuild`.
43 changes: 19 additions & 24 deletions docs/PUBLISHING_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,23 +52,24 @@ Update `CHANGELOG.md` with a dated release entry before publishing.

### 3. Test docs.rs Build Locally

docs.rs runs in a **network-isolated container** - it cannot download dependencies at build time. Our `build.rs` detects `DOCS_RS=1` and generates stub bindings instead of compiling whisper.cpp.
docs.rs runs in a **network-isolated container**, so it cannot download whisper.cpp at build time. The sys crate package therefore ships only whisper.cpp's public headers (`whisper.cpp/include/*.h`, `whisper.cpp/ggml/include/*.h`) plus whisper.cpp's `LICENSE`. When `build.rs` sees `DOCS_RS=1` it skips compiling whisper.cpp and runs the normal bindgen step against those headers, so docs.rs gets exact bindings with nothing to maintain by hand. The docs.rs image includes libclang (`clang`, `libclang-dev` in [crates-build-env](https://github.com/rust-lang/crates-build-env)).

**Test the stub bindings work:**
Regular builds of the published crate ignore the packaged headers: `build.rs` only uses the bundled `whisper.cpp/` directory when it contains the full source tree (`CMakeLists.txt` and `src/whisper.cpp`), otherwise it downloads the pinned commit.

```bash
# Clean and rebuild with DOCS_RS simulation
export DOCS_RS=1
cargo clean -p whisper-cpp-plus-sys
cargo check -p whisper-cpp-plus
**Simulate docs.rs against the packaged crate** (headers only, no network):

# Test docs generation
cargo doc -p whisper-cpp-plus --no-deps
```bash
cargo package -p whisper-cpp-plus-sys --no-verify
mkdir -p /tmp/wcp-docsrs && tar -xzf target/package/whisper-cpp-plus-sys-X.Y.Z.crate -C /tmp/wcp-docsrs
DOCS_RS=1 cargo doc --offline --no-deps \
--manifest-path /tmp/wcp-docsrs/whisper-cpp-plus-sys-X.Y.Z/Cargo.toml \
--target-dir /tmp/wcp-docsrs/target

# The high-level crate, as docs.rs builds it (all features):
DOCS_RS=1 cargo doc --offline --no-deps -p whisper-cpp-plus --all-features --target-dir /tmp/wcp-docsrs/target
```

If this fails, the stub bindings in `whisper-cpp-plus-sys/build.rs` (`generate_stub_bindings()`) need updating to include missing FFI symbols.

Unset `DOCS_RS` before running normal tests again.
On Windows, GNU tar needs `--force-local` for paths with a drive letter. Use a separate `--target-dir` so the `DOCS_RS` build script run doesn't invalidate your normal build cache.

### 4. Run Tests

Expand Down Expand Up @@ -146,24 +147,18 @@ After publishing, monitor the docs.rs build:

1. Check build queue: https://docs.rs/releases/queue
2. View build status: https://docs.rs/crate/whisper-cpp-plus/VERSION/builds
3. If build fails, check logs and fix stub bindings
3. If build fails, check the logs

### Common docs.rs Failures

| Error | Cause | Fix |
|-------|-------|-----|
| DNS resolution failed | Network access attempted | Ensure `DOCS_RS` check in build.rs |
| Cannot find function X | Missing stub binding | Add function to `generate_stub_bindings()` |
| Type mismatch | Stub signature wrong | Match stub to actual usage in high-level crate |
| Inner attribute not permitted | `#![allow(...)]` in included file | Remove inner attrs from stub bindings |

## Stub Bindings Maintenance

When adding new FFI functions to the high-level crate, also add stubs:
| DNS resolution failed | Network access attempted | Ensure the `DOCS_RS` early return in `build.rs` runs before any download or CMake step |
| `whisper.h not found` | Headers missing from the package | Check the `include` list in `whisper-cpp-plus-sys/Cargo.toml` and `cargo package -p whisper-cpp-plus-sys --list` |
| `'<header>.h' file not found` | A packaged header includes a file outside the packaged directories | Add the directory to the sys crate's `include` list |
| Unable to find libclang | docs.rs image changed | Check [crates-build-env](https://github.com/rust-lang/crates-build-env) and open an issue there |

1. Add function to `generate_stub_bindings()` in `whisper-cpp-plus-sys/build.rs`
2. Match the signature to how the high-level code calls it
3. Test with `DOCS_RS=1 cargo check -p whisper-cpp-plus`
New FFI functions need no docs.rs-specific work: bindings are generated from the same headers everywhere.

## Yanking Bad Releases

Expand Down
12 changes: 11 additions & 1 deletion whisper-cpp-plus-sys/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,17 @@ license.workspace = true
repository.workspace = true
description = "Low-level FFI bindings for whisper.cpp"
readme = "README.md"
exclude = ["whisper.cpp"]
# Only whisper.cpp's public headers (and its license) are packaged, so docs.rs, which has no
# network access, can generate exact bindings. Regular builds download the full pinned source.
include = [
"/build.rs",
"/cuda_detect.rs",
"/src/**",
"/README.md",
"/whisper.cpp/LICENSE",
"/whisper.cpp/include/*.h",
"/whisper.cpp/ggml/include/*.h",
]

[dependencies]

Expand Down
2 changes: 1 addition & 1 deletion whisper-cpp-plus-sys/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# whisper-cpp-plus-sys

> **Pinned to whisper.cpp 1.9.4-dev** (fork: [`rmorse/whisper.cpp`](https://github.com/rmorse/whisper.cpp), branch: `stream-pcm`, commit [`de8fb5fd`](https://github.com/rmorse/whisper.cpp/commit/de8fb5fda8b25837a2ba0034c8c24223a6fd6c6c), based on upstream `ggml-org/whisper.cpp` `master` after `v1.9.3`)
> **Pinned to whisper.cpp post-v1.9.4** (fork: [`rmorse/whisper.cpp`](https://github.com/rmorse/whisper.cpp), branch: `stream-pcm`, commit [`de8fb5fd`](https://github.com/rmorse/whisper.cpp/commit/de8fb5fda8b25837a2ba0034c8c24223a6fd6c6c), based on upstream `ggml-org/whisper.cpp` `master` after the `v1.9.4` release)

Low-level FFI bindings to [whisper.cpp](https://github.com/ggerganov/whisper.cpp) for Rust.

Expand Down
Loading
Loading