Skip to content

Commit 041a480

Browse files
committed
feat(jev): integrate native evaluation into Agent providers
1 parent bbf6c41 commit 041a480

43 files changed

Lines changed: 865 additions & 1552 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎apps/docs/components/ui/icon-mapping.ts‎

Lines changed: 0 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -286,7 +286,6 @@ import {
286286
TTSIcon,
287287
TwilioIcon,
288288
TypeformIcon,
289-
TypeSafeIcon,
290289
UpstashIcon,
291290
UptimeRobotIcon,
292291
VantaIcon,
@@ -475,7 +474,6 @@ export const blockTypeToIconMap: Record<string, IconComponent> = {
475474
instantly: InstantlyIcon,
476475
intercom: IntercomIcon,
477476
intercom_v2: IntercomIcon,
478-
jev: TypeSafeIcon,
479477
jina: JinaAIIcon,
480478
jira: JiraIcon,
481479
jira_service_management: JiraServiceManagementIcon,

‎apps/docs/content/docs/integrations/jev.mdx‎

Lines changed: 0 additions & 121 deletions
This file was deleted.

‎apps/docs/content/docs/integrations/meta.json‎

Lines changed: 0 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -131,7 +131,6 @@
131131
"infisical",
132132
"instantly",
133133
"intercom",
134-
"jev",
135134
"jina",
136135
"jira",
137136
"jira_service_management",

‎apps/docs/content/docs/workflows/blocks/agent.mdx‎

Lines changed: 25 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -31,6 +31,30 @@ For a custom cloud deployment, enter its provider prefix and model ID: `azure/my
3131

3232
Ollama Cloud, OpenRouter, Fireworks, Together AI, Baseten, Ollama, vLLM, and LiteLLM load their available models from the configured provider. New models appear through that discovery without a Sim catalog release. You can also enter a namespaced ID directly, such as `ollama-cloud/deepseek-v4.1-flash`, `openrouter/provider/model`, or `ollama/my-local-model`. Provider prefixes are case-insensitive; the model ID after the prefix keeps its original casing.
3333

34+
### Jev evaluation models
35+
36+
Select `jev-1.13.0`, `jev-latest`, or `jev-preview` from TypeSafe in the Agent model selector, then enter your TypeSafe API key. These models use **State** and **Questions** in place of conversational messages. State accepts text or a reference to a JSON object or array. Questions is a JSON object keyed by the answer names you want:
37+
38+
```json
39+
{
40+
"route": {
41+
"type": "choice",
42+
"instructions": "Which team should handle this request?",
43+
"criteria": { "billing": "Payments and invoices", "support": "Product issues" }
44+
},
45+
"urgency": {
46+
"type": "score",
47+
"instructions": "How urgent is the request?",
48+
"criteria": ["Routine", "Soon", "Immediate"]
49+
},
50+
"resolved": { "type": "noul", "instructions": "Has the request been resolved?" }
51+
}
52+
```
53+
54+
Read results from `<agent.answers>`. Each Choice answer includes `choice`, `probabilities`, and `confidence`; each Score answer includes `score`, `legend`, `probabilities`, and `confidence`; each Noul answer includes `noul`, a probability from 0 to 1. `content` contains the same answers as JSON text, and the standard model, token, timing, and cost outputs remain available. Use a Condition block to route on these results.
55+
56+
Jev evaluates the supplied state in one request. Chat messages, files, tools, skills, conversation memory, response-format schemas, and chat model fallbacks are hidden for these models. Saved settings return when you switch back to a chat model. Bring your own key; Sim does not provide hosted Jev credits. TypeSafe documents a 64,000-token total request limit and a 32,000-token limit for state plus the longest question. See [TypeSafe's model documentation](https://docs.typesafe.ai/models) and [question formats](https://docs.typesafe.ai/api).
57+
3458
### Files
3559

3660
Files for the model to read: images for a vision-capable model, or documents for text. Upload them on the block, or pass a file from an earlier block, such as an upload trigger or an [API](/workflows/blocks/api) response, with a connection tag.
@@ -104,7 +128,7 @@ Some settings live under advanced, or appear only for models that support them:
104128
- **Max output tokens.** Caps the response length. Defaults to the model's full limit.
105129
- **Reasoning effort / Thinking level.** For models with extended reasoning, how much the model thinks before answering. Higher is more thorough but slower and costs more tokens.
106130
- **Prompt caching.** For Anthropic Claude models, reuses the system prompt and tool definitions between runs instead of re-reading them every time. Cached input costs a tenth of the normal rate, but writing the cache costs 1.25x, so leave it off for one-off runs and turn it on when the same agent runs repeatedly. The cache covers a prefix only if it reaches 1,024 tokens (2,048 on Haiku) — below that Anthropic ignores it and nothing changes. Entries expire after five minutes of no use.
107-
- **API key.** Your key for the chosen provider. Hidden on hosted Sim, which supplies one.
131+
- **API key.** Your key for the chosen provider. Hidden when hosted Sim supplies a key for the selected model. Jev requires your own TypeSafe key.
108132
- **Fallback models.** An ordered list of up to five models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, once each, and `<agent.model>` reports the model that answered. On hosted Sim, hosted models use your workspace's BYOK or platform credentials; local and self-hosted installations may still require a key. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. A stored row key stops applying when its key field is hidden. Providers that require family-specific credentials, such as Vertex, can only be fallbacks for a selected model of the same family. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it; leave it empty and the provider's default applies.
109133
- **Retry on fail.** Retries the selected model after a failure, up to a maximum number of tries with a wait between them. When its tries run out, the fallback models are tried in order, once each, with no wait before the first of them. A fallback is never retried. See [Retries and fallbacks](#retries-and-fallbacks) for how recorded tool results are reused and when a tool can execute again.
110134

Lines changed: 72 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,72 @@
1+
import { describe, expect, it, vi } from 'vitest'
2+
import { evaluateSubBlockCondition } from '@/lib/workflows/subblocks/visibility'
3+
import { AgentBlock } from '@/blocks/blocks/agent'
4+
import { getAgentModelOptions, getModelOptions } from '@/blocks/utils'
5+
import { getBaseModelProviders } from '@/providers/models'
6+
import { Serializer } from '@/serializer'
7+
import { useProvidersStore } from '@/stores/providers/store'
8+
import type { BlockState } from '@/stores/workflows/workflow/types'
9+
10+
vi.mock('@/blocks', async () => {
11+
const { AgentBlock } = await import('@/blocks/blocks/agent')
12+
return { getBlock: () => AgentBlock }
13+
})
14+
15+
describe('Agent evaluation configuration', () => {
16+
it.each(['jev-1.13.0', 'jev-latest', 'jev-preview'])(
17+
'shows native fields and credentials for %s',
18+
(model) => {
19+
const visible = AgentBlock.subBlocks
20+
.filter((field) => evaluateSubBlockCondition(field.condition, { model }))
21+
.map((field) => field.id)
22+
expect(visible).toEqual(['model', 'evaluationState', 'evaluationQuestions', 'apiKey'])
23+
}
24+
)
25+
26+
it('keeps evaluation inputs configurable for a model reference', () => {
27+
for (const field of AgentBlock.subBlocks.filter((field) => field.id.startsWith('evaluation'))) {
28+
expect(evaluateSubBlockCondition(field.condition, { model: '<start.model>' })).toBe(true)
29+
}
30+
})
31+
32+
it('shows Jev only in the model picker that supports evaluation inputs', () => {
33+
useProvidersStore.getState().setProviderModels('base', Object.keys(getBaseModelProviders()))
34+
expect(getAgentModelOptions().map((option) => option.id)).toContain('jev-1.13.0')
35+
expect(getModelOptions().map((option) => option.id)).not.toContain('jev-1.13.0')
36+
})
37+
38+
it.each([false, true])(
39+
'serializes native fields without requiring messages, advanced=%s',
40+
(advancedMode) => {
41+
const values = {
42+
model: 'jev-1.13.0',
43+
apiKey: '{{TYPESAFE_API_KEY}}',
44+
evaluationState: '42',
45+
evaluationQuestions: '{"passed":{"type":"noul","instructions":"Did it pass?"}}',
46+
messages: JSON.stringify([{ role: 'user', content: 'Old chat prompt' }]),
47+
}
48+
const block: BlockState = {
49+
id: 'agent-test',
50+
type: 'agent',
51+
name: 'Evaluator',
52+
position: { x: 0, y: 0 },
53+
enabled: true,
54+
advancedMode,
55+
outputs: {},
56+
subBlocks: Object.fromEntries(
57+
Object.entries(values).map(([id, value]) => [
58+
id,
59+
{ id, value, type: AgentBlock.subBlocks.find((field) => field.id === id)!.type },
60+
])
61+
),
62+
}
63+
const result = new Serializer().serializeWorkflow({ [block.id]: block }, [], {}, {}, true)
64+
expect(result.blocks[0].config.tool).toBe('typesafe')
65+
expect(result.blocks[0].config.params).toMatchObject({
66+
evaluationState: '42',
67+
evaluationQuestions: values.evaluationQuestions,
68+
})
69+
expect(result.blocks[0].config.params).not.toHaveProperty('messages')
70+
}
71+
)
72+
})

0 commit comments

Comments
 (0)