Skip to content

feat(llm)!: add native Perplexity Agent API support - #62

Open
Johnson-f wants to merge 2 commits into
tinyhumansai:mainfrom
Johnson-f:feat/perplexity-agent
Open

Johnson-f wants to merge 2 commits into
tinyhumansai:mainfrom
Johnson-f:feat/perplexity-agent

Conversation

@Johnson-f

@Johnson-f Johnson-f commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add a native Perplexity Agent API provider so callers can use one explicit key with a provider-qualified model, fallback models, or a preset. Preserve rich output and deliver real incremental stream events, with explicit background recovery and generated-file access.

Related issue

None.

API or behavior changes

  • Add PerplexityModel, typed configuration/options, and ProviderKind::Perplexity. Use /v1/agent, with an explicit /v1/responses creation alias.
  • Add normalized ordered output, annotations, hosted-tool results, execution status, progress, and exact reported costs. Hosted tools remain remote; application functions return to the caller.
  • Preserve custom-function IDs, original arguments, and thought signatures through explicit replay. Include a runnable Google function-replay example.
  • Add background submit/retrieve/reconnect/cancel and response-scoped file listing/download. Bound requests, reads, stream events, timeouts, and retries; do not automatically replay ambiguous creation failures.
  • Rust source compatibility: public response/tool structs gain fields and public enums gain variants. Struct literals and exhaustive matches need updates; existing serialized records remain readable through defaults. See the migration guide. No release version is changed here.

Provider limitation observed live

Direct HTTP probes independently reproduced HTTP 400 when previous_response_id references a response ending in a custom function call, for both Google and OpenAI. Even plain user follow-ups to those parents failed; completed assistant-text parents worked. Explicit full call/result replay succeeded. The adapter preserves that error without silently resubmitting or changing strategies. The exact internal provider cause is unconfirmed.

Cancellation was acknowledged as cancelling, followed by incomplete; the adapter preserves the reported status rather than treating acknowledgement as confirmed cancellation.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo build --all-targets --all-features
  • cargo test --all-features — 1,082 passing unit, integration, and documentation tests
  • RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features
  • git diff --cached --check

Live verification covered Google /agent, OpenAI /responses, preset web search with source/citation-marker preservation, incremental text, signed function calls/replay, ordinary conversation continuation, background reconnect/retrieval, cancellation acknowledgement, and CSV generation/listing/download. cargo run -p tinyinference-llm --example perplexity_function_replay completed successfully with explicitly supplied credentials.

Rust 1.88 MSRV, default-feature-only tests, cargo-deny, and coverage measurement were not run locally. The corresponding extra toolchain/tools are not installed locally.

Tests

Offline fixtures cover request validation, known and unknown output, exact decimal cost rounding, split UTF-8/SSE frames, provisional null content/empty model fields observed live, function argument/signature preservation, authoritative snapshots, partial failures, size/deadline bounds, retry classification, credential redaction, scoped recovery handles, file access, and the public network guard.

Remote MCP/connector integrations and every model/hosted-tool combination have not been tested live.

Documentation

  • README: native provider overview.
  • docs/migrations/perplexity-agent.md: contracts, migration changes, retries, storage, and live limitations.
  • crates/tinyinference-llm/examples/perplexity_function_replay.rs: explicit signed call/result replay.

Checklist

  • The change is focused on one logical change
  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

Review follow-up

Reviewed all ten inline threads (including the two duplicated lane findings) and both additional concerns in the test-lane summary against the full source. The automated reviewer reported unavailable code retrieval; several findings were based on incomplete context.

Finding Disposition and evidence
Execution progress lost during reconstruction Fixed. Preserve nonempty progress in the execution snapshot and only backfill an empty snapshot. Regression covers both snapshot and separate-event progress.
Hosted options not validated by apply_to Fixed early validation by reusing the existing hosted-tool encoder/validator. Invalid image/fetch limits now fail without mutating the request; they were already rejected before HTTP send.
Forced hosted tool missing from declarations Fixed for explicit models/fallback chains. Presets retain server-owned tool defaults, so preset-supplied tools remain usable. Tests cover both paths; preset tools merge.
Creation alias vs lifecycle endpoints Intentional: responses_alias controls creation only, with canonical Agent recovery routes. Expanded the submit/poll/cancel test to exercise alias creation followed by /agent/{id} recovery.
Empty generic default Perplexity model Intentional explicit selection. Native constructors require a model/fallback/preset; OpenAiModel::from_spec rejects Perplexity. Clarified the ProviderSpec.model rustdoc. A Sonar default would select a different API contract.
Deferred operations bypass network guard Not reproduced. All HTTP operations call Transport::send, which checks the guard. The isolated public integration test exercises invoke, stream, submit, retrieve, resume, cancel, list, download, and deferred polling.
ToolMessage literals do not compile Not reproduced. In-repository literals were updated in the original commit; the baseline PR passed Rust stable and Rust 1.88 CI. Downstream literal/exhaustive-match changes are deliberately marked breaking and documented.
Populate OpenAI rich output / Anthropic full partial_response Outside this feature's accepted scope. Existing adapters keep their prior behavior and initialize the new optional fields; this PR does not migrate their output mapping. Perplexity populates the full rich/partial response contracts.
Integration test leaks the network flag Not reproduced. tests/perplexity_api.rs is a separate Cargo integration-test binary containing one test. Other integration files and unit tests execute in different processes. No concurrent unrelated case shares this process.
Slow send/upload bypasses the total deadline Not reproduced. timeout_at(deadline, work) wraps the entire send/retry future. Added a paused-clock regression proving a pending transport send is cancelled at the request deadline.
Reconnect discards replayed activity Intentional cursor semantics: emit only events after starting_after, with a complete authoritative terminal response. The existing reconnect regression verifies old events are skipped, new deltas survive, and the final answer is complete.

All five required local checks passed again after these changes (1,082 tests). No live paid calls were needed for this follow-up. Review threads have not been manually marked resolved.

@tinysweeper

tinysweeper Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 12 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Changes requested
Priority: high
Reviewed head: 4ace3d32a466
Updated: 1791213387 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 26 Active findings 5
Tests 13 Noted findings 0
Documentation 2 Resolved findings 141
Configuration 1 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • medium · critique · Require forced hosted tools to be declared — A request such as `PerplexitySelection::preset("low")` with `Hosted("web_search")` and no `options.tools` passes this check solely because it uses a preset. The builder cannot know (crates/tinyinference\-llm/src/providers/perplexity/request\.rs:283)
  • high · security · Provide a usable default Perplexity model — `ProviderSpec::for_kind(ProviderKind::Perplexity)` returns an empty model. Callers that construct a provider from the generic `ProviderSpec` and do not supply an explicit selection (crates/tinyinference\-llm/src/providers/types\.rs:256)
  • medium · security · Require forced hosted tools to be declared — When the selection is a preset, this condition skips the check that the forced hosted tool appears in the outgoing `tools` array. A caller can therefore force any allowlisted hoste (crates/tinyinference\-llm/src/providers/perplexity/request\.rs:283)
  • medium · security · Reject explicit tools that shadow system tools — `functions` is populated from `system_tools` first, but the duplicate check only tracks names from `request.tools`. An explicit tool can therefore reuse a reconstructed system tool (crates/tinyinference\-llm/src/providers/perplexity/request\.rs:232)
  • medium · description · Restore the network guard after the test — This finding from the previous review is unchanged. `deny_network_models()` flips the process-wide guard (`network_models_denied()` reads it back) and the test never calls `allow_n (\(pull request description\))

Resolved this pass

  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • critical — Update all ToolMessage literals for the new required field
  • medium — Preserve execution progress during reconstruction
  • critical — Preserve compatibility for existing ToolMessage literals
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Validate every hosted tool before serialization
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Critical — Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Validate every hosted tool before serialization
  • Require forced hosted tools to be declared
  • Populate normalized Responses output and execution metadata
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve compatibility for existing ToolMessage literals
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve compatibility for existing ToolMessage literals
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve compatibility for existing ToolMessage literals
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Restore the network guard after the test
  • Preserve compatibility for existing ToolMessage literals
  • Preserve execution progress during reconstruction
  • Use the configured endpoint when retrieving responses
  • Provide a valid default Perplexity model
  • Preserve execution progress in reconstructed responses
  • Guard every deferred request against network denial
  • Validate every hosted tool before serialization
  • Populate the partial response on stream failure
  • Require forced hosted tools to be declared
  • Update all ToolMessage literals for the new required field
  • Populate normalized Responses output and execution metadata
  • Preserve compatibility for existing ToolMessage literals
  • Preserve execution progress during reconstruction

Before merge

  • Address Provide a usable default Perplexity model (crates/tinyinference\-llm/src/providers/types\.rs).

How this fits together

flowchart LR
  n0["...minal_stream_item_disarms_its_abort_guard<br/>changed"]:::changed
  n1["new"]:::impacted
  n2["ModelResponse"]:::impacted
  n3["ModelStreamItem"]:::impacted
  n4["ChatModel"]:::impacted
  n5["stream"]:::impacted
  n6["ingest"]:::impacted
  n0 -->|calls| n1
  n0 -->|tests| n1
  n0 -->|uses| n5
  n1 -->|uses| n2
  n1 -->|uses| n3
  n3 -->|uses| n2
  n4 -->|uses| n2
  n5 -->|calls| n1
  n6 -->|uses| n3
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 8 files; 2 findings. (1 already reported on an earlier push) (2 earlier finding(s) still open) (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyinference\-llm/src/providers/perplexity/request\.rs — Require forced hosted tools to be declared

security

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 8 files; 4 findings. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyinference\-llm/src/providers/types\.rs — Provide a usable default Perplexity model
  • Evidence: crates/tinyinference\-llm/src/providers/perplexity/request\.rs — Require forced hosted tools to be declared
  • Evidence: crates/tinyinference\-llm/src/providers/perplexity/request\.rs — Reject explicit tools that shadow system tools

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This revision is the second-cycle Perplexity Agent adapter: it fills in the modules (`request.rs`, `config.rs`, `lifecycle.rs`, `stream.rs`, `transport.rs`) that the previous cycle only stubbed via `mod.rs`, along with the `ModelResponse.output`/`execution` fields, `ToolCall.replay`, `ProviderKind::Perplexity`, and `StreamAccumulator` rich-output folding. Several earlier findings (test-file network-guard restoration, `ToolMessage` literal fallout, `for_kind` default model, output/execution population) are addressed by the new tests and defaults, but the provider still has the delayed-response `Progress`/`progress` mismatch, the missing endpoint propagation into `retrieve_at`, and the `CapturedModel`/`ToolMessage` field churn that the previous cycle flagged. The remaining concern is that a wide swath of the new code paths (lifecycle, streaming, request building) ship with tests that exercise the happy path but do not pin the specific invariants the doc comments assert. (12 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The incremental commits add rich-output/execution/progress collection to `StreamAccumulator` with tests, which resolves the earlier reconstruction and partial-response findings, along with the ToolMessage/ToolCall compatibility fields, endpoint-scoped handles, network-guard coverage for every deferred operation, hosted-tool pre-serialization validation, and forced-hosted-tool declaration checks. One earlier finding remains unfixed: the integration test denies network models and never restores the guard. With that one item outstanding, the change is otherwise sound and matches its description. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: \(pull request description\) — Restore the network guard after the test

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash, deepseek-v4.1-flash
  • Spend: $0.051624
  • Tokens: 953109 input · 59743 output · 62840 cached · 0 embedding
  • Continuity: summary cache chain restarted at the storage ceiling.
Head State Pass summary
c8441a1ec488 incomplete 12 active finding(s), 0 resolved finding(s) (at 1791142292)
4ace3d32a466 changes requested 5 active finding(s), 141 resolved finding(s) (at 1791213387)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 4b4e13a3-cf03-4555-8f5d-3ea3df82f637
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@Johnson-f

Copy link
Copy Markdown
Contributor Author

@senamakel pls review

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is critical.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0731 · 1,237,432 in / 56,882 out · 96,922 cached (8%)  · gpt-5.6-luna, glm-5.3-flash, , gpt-6-luna
critique:    $0.0480 · 673,614 in   / 36,421 out · 61,146 cached (9%)  · gpt-5.6-luna, glm-5.3-flash,
security:    $0.0200 · 292,907 in   / 14,956 out · 35,776 cached (12%) · gpt-5.6-luna
tests:       $0.0008 · 87,709 in    / 1,465 out  · 0 cached (0%)       · glm-5.3-flash
description: $0.0008 · 87,349 in    / 80 out     · 0 cached (0%)       · glm-5.3-flash

) -> Result<ModelResponse> {
self.validate_handle(handle)?;
let outgoing =
self.inner

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique likely

Use the configured endpoint when retrieving responses

When responses_alias is enabled, submission uses create_path() and therefore targets /responses, but all lifecycle retrieval paths here hard-code /agent. A handle created under that configuration can consequently be submitted successfully and then retrieved, resumed, cancelled, or have files accessed through the wrong endpoint. Derive the collection segment from the same configuration used for creation, or store the endpoint in the handle metadata and validate/use it consistently.

[RULE] endpoint-mismatch ·

/// Returns the default provider spec for a known provider.
pub fn for_kind(kind: ProviderKind) -> Self {
match kind {
ProviderKind::Perplexity => Self::new(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique likely

Provide a valid default Perplexity model

ProviderSpec::model is documented as the default provider model id, but the new Perplexity default sets it to an empty string. A caller using ProviderSpec::for_kind(ProviderKind::Perplexity) therefore receives a configuration with no model and may send an invalid request or fail before making a request. Use the adapter's documented default model (for example, a supported sonar model), or verify that the adapter intentionally fills this value before the request; the adapter implementation was not included in the reviewed context.

[RULE] invalid-default ·

};
Ok(ModelResponse {
output: self.output.into_values().collect(),
execution: self.execution.map(|mut execution| {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique confident

Preserve execution progress in reconstructed responses

When a stream has an Execution event whose ModelExecution.progress is already populated but has no separate Progress events, this assignment replaces that progress with the empty accumulator vector. The completed-response path explicitly preserves existing progress and only backfills when it is empty, so the no-Completed path should apply the same rule.


Additional security observation

priority medium confident

Preserve execution progress during reconstruction

[RULE] preserve-stream-metadata

When no Completed item is received, this unconditionally replaces progress already present on execution with self.progress. If progress was supplied in the execution event but no separate Progress events were emitted, the reconstructed response silently loses it. Only backfill progress when the execution's existing progress is empty, matching the handling used for completed responses.

Suggested change for this observation (reference only)

execution: self.execution.map(|mut execution| {
                if execution.progress.is_empty() {
                    execution.progress = self.progress;
                }
                execution
            }),

[RULE] preserve-stream-state ·

let outgoing =
self.inner
.request(reqwest::Method::GET, &path(&["agent", &handle.id])?, None)?;
let incoming = self.inner.send(outgoing, Operation::Read, deadline).await?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique confident

Guard every deferred request against network denial

submit_background checks ensure_network_models_allowed, but retrieval, resume, cancellation, file listing, and file download issue network requests without that check. After a caller creates a handle, calling deny_network_models() does not prevent these later requests, contrary to the repository contract that every request-issuing path calls the guard. Add the guard to each public deferred-operation entry point (or otherwise ensure every individual request is guarded).

[RULE] network-guard ·

/// # Errors
/// Returns a serialization error if options cannot be encoded.
pub fn apply_to(&self, request: &mut ModelRequest) -> Result<()> {
super::request::validate_options_before_serialization(self)?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique confident

Validate every hosted tool before serialization

The validation invoked here only checks PerplexityTool::WebSearch; ImageSearch.max_results and FetchUrl.max_urls are not validated. For example, ImageSearch { max_results: Some(0), .. } and FetchUrl { max_urls: Some(11) } are serialized successfully even though their documented ranges are 1–30 and 1–10. Reject these values before writing provider_options, otherwise callers can reach the provider with invalid requests and receive avoidable request failures.

[RULE] missing-validation ·

crate::failure::classify_provider_failure(None, code.as_deref(), &message)
.is_retryable();
return Err(Error::Provider(Box::new(ProviderError {
partial_response: None,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique likely

Populate the partial response on stream failure

When content has already been received and the stream later fails, provider_failure consumes the accumulator into a partial AssistantMessage but leaves ProviderError::partial_response as None. Callers therefore cannot access the accumulated usage, raw response, finish state, or the normalized full ModelResponse through the documented partial_response field; they only receive the legacy partial_message. Build the accumulated response once and assign it to partial_response while retaining its message for partial_message.

[RULE] preserve-partial-response ·

if !functions.contains_key(name) {
return Err(invalid("selected function is not declared"));
}
} else if let Some(kind) = choice.get("type").and_then(Value::as_str)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique likely

Require forced hosted tools to be declared

A PerplexityToolChoice::Hosted("web_search") (or another allowed hosted kind) passes this validation even when options.tools contains no corresponding hosted tool. The resulting payload includes a forced tool_choice without declaring that tool, which the provider cannot honor and may reject. Validate that the selected hosted kind is present in the encoded hosted tools before inserting the choice.

[RULE] validate-tool-choice ·

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved — the review agent found this finding fixed in the new code, as of 4ace3d3.

If this is wrong, reopen the conversation and say so; the finding will be re-raised on the next push if it still reproduces.

pub struct ToolMessage {
/// Name and signed-call context for stateful native function results.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub call_context: Option<crate::tool::ToolResultContext>,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority critical security confident

Update all ToolMessage literals for the new required field

Adding a field to this public struct breaks existing struct literals that construct ToolMessage without call_context (including production and test callers elsewhere in the crate). Rust requires every struct literal to specify the new field, so the crate will fail to compile as written. Update all callers in the same change, or redesign this API so adding context does not require every literal to change.


Additional critique observation

priority critical confident

Preserve compatibility for existing ToolMessage literals

[RULE] breaking-api

Adding a field to this public struct requires every existing struct literal to initialize it. The repository still has ToolMessage { ... } literals that do not provide call_context, including production code, so the workspace will fail to compile; downstream callers using the public struct literal syntax break as well. Update all construction sites (using call_context: None where appropriate), or provide a compatibility-preserving construction/API strategy before merging.

[RULE] breaking-public-struct-change ·

_ => Some("stop".to_string()),
};
ModelResponse {
output: Vec::new(),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security confident

Populate normalized Responses output and execution metadata

The OpenAI Responses endpoint is not a legacy provider, yet this parser unconditionally reports an empty ordered output and no execution metadata. Downstream callers therefore lose the provider's ordered reasoning, text, tool-call, status, identity, and detailed usage information despite ModelResponse explicitly exposing those normalized fields. Map the parsed Responses items and response metadata into ModelOutputItem and ModelExecution instead of defaulting them away.

[RULE] preserve-response-metadata ·

});
let mut handle = model.response_handle(&snapshot).unwrap();
handle.kind = Some("background".into());
deny_network_models();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security likely

Restore the network guard after the test

This test changes a process-wide network policy and never restores it. Because integration tests can share the same process and run concurrently, the guard can leak into unrelated tests, causing nondeterministic failures or masking network-dependent behavior. Use a scoped guard or restore the prior state before the test exits.

[RULE] global-state-leak ·

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0516 · 953,109 in / 59,743 out · 62,840 cached (7%) · gpt-5.6-luna, glm-5.3-flash, deepseek-v4.1-flash
critique:    $0.0227 · 373,734 in / 23,986 out · 34,580 cached (9%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0196 · 307,923 in / 23,321 out · 26,852 cached (9%) · gpt-5.6-luna
tests:       $0.0000 · 85,561 in  / 229 out    · 1,408 cached (2%)  · deepseek-v4.1-flash
description: $0.0091 · 89,644 in  / 9,460 out  · 0 cached (0%)      · glm-5.3-flash

/// Returns the default provider spec for a known provider.
pub fn for_kind(kind: ProviderKind) -> Self {
match kind {
ProviderKind::Perplexity => Self::new(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high security confident

Provide a usable default Perplexity model

ProviderSpec::for_kind(ProviderKind::Perplexity) returns an empty model. Callers that construct a provider from the generic ProviderSpec and do not supply an explicit selection will therefore pass an invalid model to the Perplexity adapter, despite this being the advertised default provider configuration. Use a supported Perplexity model as the default, or make the generic provider-construction path require and validate an explicit selection instead of returning an apparently usable spec.

[RULE] invalid-default ·

{
return Err(invalid("unsupported forced hosted tool"));
}
if !matches!(selection, PerplexitySelection::Preset { .. })

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security confident

Require forced hosted tools to be declared

When the selection is a preset, this condition skips the check that the forced hosted tool appears in the outgoing tools array. A caller can therefore force any allowlisted hosted tool type without declaring it, producing an inconsistent request and potentially invoking provider-managed capabilities that were not part of the configured tool set. Require the selected hosted tool to be declared regardless of whether the model selection is a preset.


Additional critique observation

priority medium confident

Require forced hosted tools to be declared

[RULE] undeclared-tool-choice

A request such as PerplexitySelection::preset("low") with Hosted("web_search") and no options.tools passes this check solely because it uses a preset. The builder cannot know that an arbitrary or future provider preset actually exposes the selected hosted tool, so the request can reach the provider with a forced tool choice that is not present in the declared tool list and be rejected or behave inconsistently. Validate the hosted choice against the declared tools for presets too, or explicitly resolve and validate the preset's tool loadout before accepting it.

Suggested change for this observation (reference only)

if !body
                .get("tools")
                .and_then(Value::as_array)
                .is_some_and(|tools| {

[RULE] undeclared-tool-choice ·

functions.insert(tool.name.clone(), tool);
}
let mut explicit_names = HashSet::new();
for tool in &request.tools {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security confident

Reject explicit tools that shadow system tools

functions is populated from system_tools first, but the duplicate check only tracks names from request.tools. An explicit tool can therefore reuse a reconstructed system tool's name and overwrite its schema in the outgoing request, changing the effective tool set and potentially routing a call to the wrong implementation. Reject names already present in functions before inserting explicit tools.

[RULE] tool-name-collision ·

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant