Skip to content

feat: add moss knowledge retrieval extension and example - #2351

Open
rafaelbeckel wants to merge 1 commit into
TEN-framework:mainfrom
rafaelbeckel:feat/moss-knowledge-retrieval
Open

rafaelbeckel wants to merge 1 commit into
TEN-framework:mainfrom
rafaelbeckel:feat/moss-knowledge-retrieval

Conversation

@rafaelbeckel

Copy link
Copy Markdown

Summary

Adds Moss knowledge retrieval to TEN in two parts:

  • moss_tool_python: a search_knowledge_base tool for any TEN graph. It loads a Moss index into memory when the agent starts and runs hybrid (semantic and keyword) search locally, so a search takes a few milliseconds and makes no network request.
  • examples/voice-assistant-with-moss: the voice-assistant example with two graphs. moss_ambient searches on every final user turn and adds the passages to that turn, so the LLM answers in one call. moss_tool lets the LLM call the tool when it needs to. Ambient retrieval is a context_tool property on main_control (about 20 lines); the rest of the example is copied from voice-assistant.

Why

  • Who it is for. About 80% of Moss customers build voice agents on LiveKit and know its rooms model. TEN's rooms on Agora RTC give them a familiar way in, and a ready example lets them try the integration and evaluate Agora for their own deployments.
  • What it adds. Validation showed no significant technical advantage over running Moss behind the Custom LLM path of Agora's Conversational AI Engine. The gain is familiarity and speed: teams that build on TEN get document grounding without writing an integration, and can reach production in days to weeks.

Architecture and configuration

Both modes, with diagrams, are in the example README. The node settings (project, key, index, top_k) are in the extension README. Credentials come from MOSS_PROJECT_ID, MOSS_PROJECT_KEY and MOSS_INDEX_NAME, added to .env.example.

Dependency

moss>=1.15.0 from PyPI, which adds one native wheel (inferedge-moss-core) for Linux x86_64 (glibc 2.35+), Linux aarch64, macOS arm64 and Windows. Both are source-available under PolyForm Shield 1.0.0: free for any use, including commercial, except building a product that competes with Moss. Please say if that is a problem.

Testing

Tested 2026-10-06 on this branch in TEN's dev container (ten_agent_build:0.7.14), with moss 1.15.0, a free Moss project holding the portal's 55-document support FAQ index, and the example's vendors (Deepgram, gpt-4o, ElevenLabs).

  • Voice calls in the playground answered from the index in both graphs: support hours, the 15% restocking fee on opened electronics, and order consolidation within 1 hour of ordering. In moss_ambient the LLM made no tool call; in moss_tool it called search_knowledge_base for each question.

  • Each search took 0 to 13 ms. Loading the index took about 2.9 s per session, in the background while the greeting plays.

  • From the final transcript to the first sentence sent to TTS:

    Graph Median Answers LLM calls per answer
    moss_ambient 764 ms 5 1
    moss_tool 2,132 ms 3 2
  • With a wrong project key, the extension logs that the index was not loaded and registers no tool, and the agent keeps running. The key does not appear in the logs.

  • black (line length 80) and pylint with the repo's rcfile pass for the extension.

Maintenance

Moss maintains this extension and the example: testing, documentation and compatibility updates, contributed proactively. We would appreciate advance notice of significant breaking changes in TEN, for example to the tool_register and tool_call messages or to main_python.

Add moss_tool_python, a search_knowledge_base tool that searches a Moss
index loaded in memory, and voice-assistant-with-moss, which runs it in
two graphs: ambient retrieval through a context_tool property on
main_control, and a plain LLM tool call.
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

Findings

  • [P1] Treat retrieved passages as untrusted input — ai_agents/agents/examples/voice-assistant-with-moss/tenapp/ten_packages/extension/main_python/extension.py:211-226
    Ambient retrieval concatenates arbitrary Moss document text directly into the same user message as the query. A document from an uploaded file or website can therefore inject instructions (for example, to ignore the system prompt or reveal conversation data), and the model has no boundary telling it that this is reference data. Delimit/quote the passages and add an explicit instruction to treat them as untrusted evidence (or use a dedicated retrieval/tool message), with a regression case containing an instruction-bearing document.

  • [P1] Bound retrieval context before retaining it in conversation history — ai_agents/agents/examples/voice-assistant-with-moss/tenapp/ten_packages/extension/main_python/agent/llm_exec.py:151-169
    _with_context() returns the passages plus the user text, and _send_to_llm() appends that entire enriched message to self.contexts on every turn. contexts has no eviction or token budget, so three retrieved passages are sent again on every subsequent turn and the request eventually exceeds the model context window (the max_memory_length graph property is not applied here). Keep retrieved context ephemeral or persist only the original query, and trim history by tokens.

  • [P2] Do not permanently disable retrieval after a transient startup failure — ai_agents/agents/ten_packages/extension/moss_tool_python/extension.py:34-47
    The extension awaits the network/index load inside on_start, catches every exception, and returns without scheduling a retry. A temporary Moss outage or slow startup therefore leaves the tool absent for the whole worker lifetime; in ambient mode _context_source stays empty. Load with a bounded timeout in a background task and retry (or expose a clear health state), while still allowing the graph to start.

  • [P2] Propagate Python dependency installation failures — ai_agents/agents/examples/voice-assistant-with-moss/tenapp/scripts/install_python_deps.sh:20-27
    The new moss dependency is installed in a loop without set -e or checking the command status. If the native wheel cannot be downloaded or is unsupported, the script prints “completed” and exits successfully, so task install appears green and the example only fails later with ModuleNotFoundError. Make each install failure terminate the script.

  • [P2] Add extension regression coverage — ai_agents/agents/ten_packages/extension/moss_tool_python/
    This new extension has no tests for successful index loading, disabled/error startup, client close, tool metadata, query arguments, or result formatting. Add mocked-client tests for those paths and an ambient retrieval test that checks prompt boundaries and history trimming.

Checks

  • python -m compileall passed for all added Python files.
  • All added JSON files parsed successfully.
  • No runtime/guarder test was run from the base worktree.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant