Repository navigation
fix(logs): close the log-once coverage gaps from the cross-PR review - #8881
waleedlatif1 wants to merge 5 commits into
Conversation
Routes the remaining execution boundaries (execute-workflow, schedule, table group cell, resume) through logFailureOnce with their execution id, so a revoked credential logs at INFO and a failure an inner boundary logged is not repeated. Collapses a non-revoked credential failure to one ERROR line at the tool boundary with the credential id, keeping the cause at WARN in the token resolver under the same request id. The child failure line carries the projected root error and the child execution id. A legacy job's usage-limit or suspended-account refusal is the author's and is no longer logged twice. The v2 execute and MCP serve responses no longer carry the internal admission codes, restoring their prior shape, and the Agent failure line names its provider and model again.
Keeps the token resolver's refresh failure at ERROR (it is the only server line on the browser token route) with the credential, provider, and redacted cause; throws a UserFailure for a legacy job's admission refusal; narrows the Agent failure details to provider and model; and passes the child execution id through unchanged.
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
There was a problem hiding this comment.
All reported issues were addressed across 19 files
Reply with feedback, questions, or to request a fix.
Turn on auto-fix | Re-trigger cubic
|
…esumes on the parent run The resume and child failure lines keep their execution, block, and resume ids outside the secret projection, which drops everything when it fails closed. Resumed runs execute under the parent's execution id, so the resume boundaries dedupe on it; a 4xx resume admission refusal is the caller's. The tool failure line names the credential selected after normalization.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
| workflowId: executionContext?.workflowId ?? undefined, | ||
| executionId: executionContext?.executionId, | ||
| blockId: typeof toolContext.blockId === 'string' ? toolContext.blockId : undefined, | ||
| credentialId: selectedCredentialId, |
There was a problem hiding this comment.
selectedCredentialId can contain a decrypted secret, not just an account ID. A direct snowflake_execute_sql call can pass credentialId: '{{TOKEN}}'. The tool's oauthCredential field is user-only, so resolveToolEnvReferences replaces the reference before this value is captured.
If credential lookup rejects that value, this catch writes the secret as credentialId, outside projectToolLogMetadata. Anyone with access to the server logs could read it. Project this field too, or log only an ID confirmed by credential lookup.
How this was verified: The public call accepts the reference, the tool resolves it from the caller's environment, and the failure log writes the resolved value without secret projection.
| if (error instanceof ResumeAdmissionError && error.statusCode < 500) { | ||
| markFailureKind(error, 'user') |
There was a problem hiding this comment.
Server failures look user-caused
Not every ResumeAdmissionError below 500 is the caller's fault. requireResumeDeploymentVersion throws it with status 409 when a saved run is missing its execution mode or deployment version, or when those values disagree. runResumeExecution checks these values from the saved snapshot and claimed log row.
The new blanket mark turns those server-owned state failures into INFO lines attributed to the user. Operators can miss broken saved state because it looks like an ordinary user mistake. Keep these failures internal and mark only refusals the caller can cause.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Summary
Closes the gaps a cross-PR review found after #8872, where author-caused failures were still logged at ERROR or logged more than once.
execute-workflow(streaming/chat API, copilot, workflow tests), the scheduled-workflow catch, the table group-cell catch, and both resume paths now go throughlogFailureOncewith their execution id, like the other surfaces.credentialId,providerId, and a redacted cause. It stays at ERROR because it is the only server line on the browser token route.toolId,workflowId,executionId,blockId, andcredentialId.credential-tokenand the token catch intools/index.ts).UserFailureand is no longer logged twice at ERROR./workflows/{id}/executeand MCP serve no longer return the internal admission codesUSAGE_LIMIT_EXCEEDEDandACCOUNT_SUSPENDEDindetails.code/ error data.ACCOUNT_SUSPENDEDis not inFORBIDDEN_DETAIL_CODES, the closed set the v2 contract allows on a 403.USAGE_LIMIT_EXCEEDED.Kept at ERROR: unattributed faults at every boundary, and the token resolver's refresh failure (it holds the cause).
What to watch in CloudWatch after deploy
failureKindfield for author-caused failures, and mostly disappear when an inner boundary already logged the failure.credentialIdandproviderId.details.codefor these admission codes.Type of Change
Testing
execute-workflow: revoked credential and internal fault are both logged once for the execution (probed throughlogFailureOnce).details.code.bunx biome checkon changed files. Full lint, type-check,check:audits, and suites run in CI.Checklist
test-auditauthoring gate)