Skip to content

Commit 53160eb

Browse files
committed
feat(memory): bound durable context and retrieve retained tool detail
1 parent 94da520 commit 53160eb

96 files changed

Lines changed: 34987 additions & 573 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎apps/docs/content/docs/workflows/blocks/agent.mdx‎

Lines changed: 7 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -64,15 +64,19 @@ When durable tool history is enabled for your workspace, memory also keeps the a
6464

6565
Tool calls and their results are selected together, including parallel calls. They do not each use another slot in a message-count window, but their arguments and results still consume input tokens. Use a token window when the amount of recalled context matters more than the number of messages. Windowing changes what the model receives, not what is stored.
6666

67-
Large results are retained separately with a bounded preview in history. A preview can omit details: storing the full result does not mean the model can see it, and there is no automatic retrieval of omitted content or conversation summarization. Sim also budgets recalled history for the selected primary or fallback model, leaving room for instructions, tools, attachments, and output. These limits do not cap the total tokens used by a run; new tool results can grow context during its tool loop.
67+
Large results are retained separately, with their first 8,000 characters in model context and a notice when the result is truncated. The built-in `agent_memory_read` tool lets the agent search retained history and read omitted result details in small pages. It can only read the current conversation. When exact details matter, ask the agent to check the original result instead of relying on its preview.
6868

69-
The Memory API and Memory block still return plain conversation messages. Internal tool history and retry state are not added to their `data` responses.
69+
Sim normally targets up to 16,000 estimated tokens of recalled history, subject to your memory window and the model's available context. Large or difficult-to-tokenize content uses a conservative estimate. It checks the input before every model generation, including generations after tool calls and on fallback models, leaving room for instructions, tool definitions, attachments, and output. The current request and required tool exchanges stay intact. If they cannot fit, the Agent stops with a context-limit error. This limits context per generation, not the total tokens used across a run.
70+
71+
When a generation would omit older history, Sim can create a concise summary while keeping recent exchanges, including during long tool loops. Summaries can omit details and do not replace the stored conversation. Generating one uses an additional model call whose tokens and cost are included in the Agent's usage; a cached summary can be reused when its source history is unchanged. If summarization is unavailable, the Agent continues with bounded history selection.
72+
73+
The Memory API and Memory block still return plain conversation messages. Internal tool history, retry state, and cached summaries are not added to their `data` responses.
7074

7175
#### Retries and fallbacks
7276

7377
With durable tool history enabled, retries and fallback models continue the same Agent invocation using recorded tool results. For example, if a tool returns an order number and final generation fails, the fallback receives that result without calling the tool again. A new workflow execution or loop iteration is a separate invocation.
7478

75-
A call whose outcome was not recorded can execute again, including when an external action succeeded just before a failure. Use tools that safely handle repeated requests for actions that must not happen twice. If durable history is disabled or persistence is unavailable, saved-progress recovery is not guaranteed. This does not restart crashed workflows, override cancellation, or retry failures marked nonretryable. A failure after streaming output has started is not restarted on another model.
79+
A recorded terminal outcome is kept even when its stored details become unavailable; the Agent does not repeat that action merely to recover the missing details. A call whose outcome was not recorded can execute again, including when an external action succeeded just before a failure. Use tools that safely handle repeated requests for actions that must not happen twice. If durable history is disabled or persistence is unavailable, saved-progress recovery is not guaranteed. This does not restart crashed workflows, override cancellation, or retry failures marked nonretryable. A failure after streaming output has started is not restarted on another model.
7680

7781
### Response Format
7882

‎apps/sim/executor/handlers/agent/agent-handler.ts‎

Lines changed: 22 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,6 +20,8 @@ import { resolveMcpToolBinding } from '@/lib/mcp/tool-binding'
2020
import type { McpToolSchema } from '@/lib/mcp/types'
2121
import { createMcpToolId } from '@/lib/mcp/utils'
2222
import { type AgentTurnSession, openAgentTurnSession } from '@/lib/memory/agent-turn-session'
23+
import { MEMORY } from '@/lib/memory/constants'
24+
import { createAgentMemoryRetrievalTool } from '@/lib/memory/retrieval-tool'
2325
import {
2426
type AutoMediaKind,
2527
type AutoRoutingResult,
@@ -2827,6 +2829,7 @@ export class AgentBlockHandler implements BlockHandler {
28272829
config
28282830

28292831
const validMessages = this.validateMessages(messages)
2832+
const configuredHistoryTokens = Number(inputs.slidingWindowTokens)
28302833

28312834
const { blockData, blockNameMapping } = collectBlockData(ctx)
28322835

@@ -2857,6 +2860,12 @@ export class AgentBlockHandler implements BlockHandler {
28572860
userId: ctx.userId,
28582861
executionId: ctx.executionId,
28592862
stream: streaming,
2863+
memoryHistoryTokens:
2864+
inputs.memoryType === 'sliding_window_tokens'
2865+
? Number.isFinite(configuredHistoryTokens) && configuredHistoryTokens > 0
2866+
? Math.floor(configuredHistoryTokens)
2867+
: MEMORY.DEFAULT_SLIDING_WINDOW_TOKENS
2868+
: undefined,
28602869
messages: messages?.map((message) => {
28612870
const { executionId, ...providerMessage } = message
28622871
copyNativeConversationMessage(message, providerMessage)
@@ -2919,6 +2928,12 @@ export class AgentBlockHandler implements BlockHandler {
29192928
}
29202929

29212930
const { blockData, blockNameMapping } = collectBlockData(ctx)
2931+
const agentMemoryRetrieval = agentConversation?.memoryId
2932+
? createAgentMemoryRetrievalTool({
2933+
executionContext: ctx,
2934+
memoryId: agentConversation.memoryId,
2935+
})
2936+
: undefined
29222937

29232938
const response = await executeProviderRequest(
29242939
providerId,
@@ -2927,7 +2942,9 @@ export class AgentBlockHandler implements BlockHandler {
29272942
systemPrompt:
29282943
'systemPrompt' in providerRequest ? providerRequest.systemPrompt : undefined,
29292944
context: 'context' in providerRequest ? providerRequest.context : undefined,
2930-
tools: providerRequest.tools,
2945+
tools: agentMemoryRetrieval
2946+
? [...(providerRequest.tools ?? []), agentMemoryRetrieval.tool]
2947+
: providerRequest.tools,
29312948
temperature: providerRequest.temperature,
29322949
maxTokens: providerRequest.maxTokens,
29332950
apiKey: finalApiKey,
@@ -2969,6 +2986,10 @@ export class AgentBlockHandler implements BlockHandler {
29692986
resolvedSecretTraceRegistry: modelRuntimeRegistry,
29702987
executionContext: ctx,
29712988
agentConversation,
2989+
agentMemoryRetrieval,
2990+
agentMemoryContext: agentConversation
2991+
? { historyTokens: providerRequest.memoryHistoryTokens }
2992+
: undefined,
29722993
}
29732994
)
29742995

‎apps/sim/executor/handlers/agent/memory.durability.test.ts‎

Lines changed: 35 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -70,6 +70,41 @@ describe('optional Agent memory durability failures', () => {
7070
).resolves.toEqual([])
7171
})
7272

73+
it('defers rich conversation selection until the actual provider context is known', async () => {
74+
const history = [
75+
{ role: 'user', content: 'x'.repeat(140_000) },
76+
{ role: 'assistant', content: 'recent answer' },
77+
]
78+
mocks.prefix.mockResolvedValue({
79+
id: options.memoryId,
80+
storageVersion: 2,
81+
data: history,
82+
secretProvenanceVersion: 1,
83+
provenanceContentHash: hashDurableSecretProvenanceValue(history),
84+
provenanceStatus: 'exact',
85+
provenanceEntries: [],
86+
})
87+
const memory = new Memory()
88+
await expect(
89+
memory.fetchMemoryMessages(ctx, { ...inputs, model: 'gpt-4o' }, undefined, {
90+
richHistory: true,
91+
})
92+
).resolves.toEqual(history)
93+
await expect(
94+
memory.fetchMemoryMessages(
95+
ctx,
96+
{
97+
...inputs,
98+
memoryType: 'sliding_window_tokens',
99+
slidingWindowTokens: '100',
100+
model: 'gpt-4o',
101+
},
102+
undefined,
103+
{ richHistory: true }
104+
)
105+
).resolves.toEqual([history[1]])
106+
})
107+
73108
it('never rereads a replacement key after the original conversation disappears', async () => {
74109
mocks.prefix.mockResolvedValue(undefined)
75110
await expect(

‎apps/sim/executor/handlers/agent/memory.ts‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -141,7 +141,9 @@ export class Memory {
141141

142142
switch (inputs.memoryType) {
143143
case 'conversation':
144-
messages = selectConversationContextWindow(stored.messages, inputs.model, stored.groups)
144+
messages = options.richHistory
145+
? stored.messages
146+
: selectConversationContextWindow(stored.messages, inputs.model, stored.groups)
145147
break
146148

147149
case 'sliding_window': {

0 commit comments

Comments
 (0)