Repository navigation
feat(embeddings): add resilient voyage reranking support - #59
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Comment |
Tiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Incomplete Review snapshot
Completeness: Incomplete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred.
FindingsNo active actionable findings. Resolved this pass
Could not review: README.md, crates/tinyinference-embeddings/Cargo.toml, crates/tinyinference-embeddings/src/error.rs, crates/tinyinference-embeddings/src/lib.rs, crates/tinyinference-embeddings/src/rerank/mod.rs, crates/tinyinference-embeddings/src/rerank/mod_tests.rs, crates/tinyinference-embeddings/src/rerank/types.rs, crates/tinyinference-embeddings/src/rerank/voyage.rs, crates/tinyinference-embeddings/src/rerank/voyage_tests.rs, crates/tinyinference-embeddings/src/types.rs, crates/tinyinference-voice/src/postprocess.rs, crates/tinyinference-voice/src/postprocess_tests.rs, tinysweeper/tests Before merge
How this fits togetherflowchart LR
n0["cleanup_transcription<br/>changed"]:::changed
n1["llm_error_response_falls_back_to_raw_text<br/>changed"]:::changed
n2["runtime_and_service"]:::impacted
n3["spawn_mock"]:::impacted
n4["chat_response"]:::impacted
n5["degraded_service_still_runs_cleanup"]:::impacted
n6["empty_llm_response_falls_back_to_raw_text"]:::impacted
n1 -->|calls| n0
n1 -->|tests| n0
n1 -->|calls| n2
n1 -->|tests| n2
n1 -->|calls| n3
n1 -->|tests| n3
n5 -->|calls| n0
n5 -->|tests| n0
n5 -->|calls| n2
n5 -->|tests| n2
n5 -->|calls| n3
n5 -->|tests| n3
n5 -->|calls| n4
n5 -->|tests| n4
n6 -->|calls| n0
n6 -->|tests| n0
n6 -->|calls| n2
n6 -->|tests| n2
n6 -->|calls| n3
n6 -->|tests| n3
n6 -->|calls| n4
n6 -->|tests| n4
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
@senamakel pls review. |
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: README.md, crates/tinyinference-embeddings/Cargo.toml, crates/tinyinference-embeddings/src/error.rs, crates/tinyinference-embeddings/src/lib.rs, crates/tinyinference-embeddings/src/rerank/mod.rs, crates/tinyinference-embeddings/src/rerank/mod_tests.rs, crates/tinyinference-embeddings/src/rerank/types.rs, crates/tinyinference-embeddings/src/rerank/voyage.rs and 5 more.
$0.0017 · 51,922 in / 3,905 out · 0 cached (0%) · deepseek/deepseek-v4-flash
description: $0.0008 · 26,992 in / 162 out · 0 cached (0%) · deepseek/deepseek-v4-flash
|
looks great |
Summary
Add standalone Voyage reranking to
tinyinference-embeddings. Callers can rerank document texts from any search system and map validated results back to the original candidates without changing stored vectors, metadata, or similarity scores.Also make the existing voice cleanup timeout test deterministic: it now exercises unfinished inference directly instead of waiting on socket activity while Tokio time is paused.
Related issue
None.
API or behavior changes
Reranker, request/result/usage types, cancellation,VoyageReranker, andVoyageRerankConfig.rerank-3, a 30-second total deadline, bounded request/response bodies, and up to three retries for explicit 429/500/502/503/504 responses. Transport interruptions are not replayed.Error::Rerank(RerankError): downstream exhaustive matches on the existing error enum need a new arm.Validation
All passed locally against the committed contents:
cargo fmt --all -- --checkcargo clippy --all-targets --all-features -- -D warningscargo build --all-targets --all-featurescargo test --all-featuresRUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-featuresLive verification used the public Rust adapter with retries disabled: one
rerank-3request correctly selected the top two of three documents, preserved original positions, and reported 75 input tokens in 839 ms. The temporary runner was removed; credentials were not saved.Tests
Added 25 reranking tests covering request shape, result validation, usage, original-document mapping, both candidate sources, cancellation, deadlines, retry boundaries, byte limits, and credential-safe diagnostics. The embedding package passes 100 unit tests and seven documentation tests.
The updated voice timeout test checks that work is still pending before its deadline, returns the original text afterward, and drops the unfinished operation. The complete workspace suite now finishes without skipping that test.
Failure-path tests use synthetic provider responses and controlled time; live verification covers one successful request, not provider outages or billing guarantees.
Documentation
Updated README.md and the rerank module rustdoc with configuration, source-text ownership, result mapping, caller fallback, retry/billing limits, and the new error arm.
Checklist
#[allow(...)],#[ignore], or relaxed lints.envcontents in the diff or the description