[#521] Research Agent Tools - #524
Conversation
Add ask_database as a second tool for the assistant, alongside the existing semantic search_documents tool. - Reuse the SELECT-only, ALLOWED_TABLES-guarded ask_database implementation from services/tools/database.py rather than reimplementing the query guards in the assistant. - Add get_tools_schema() / make_tool_mapping() as an aggregation seam in tool_services.py so assistant_services.py no longer names individual tools; new tools are registered in one place. - Build the ask_database schema in the flattened Responses-API shape (not the nested Chat Completions shape from services/tools), and defer the database_schema_string import to call time so importing the module never triggers a DB query. - Split tool_services.py: move search_documents into search_tool.py and the agentic-loop helpers (handle_tool_calls_with_reasoning, invoke_functions_from_response) into agentic_loop.py, adding the imports each module needs. - Point importers (assistant_services.py, test_tool_services.py) at the defining module for each symbol instead of re-exporting through tool_services.py. - Document the search/SQL tool overlap risk and the import-time DB access caveat inline.
Replace the two parallel tool registries (get_tools_schema / make_tool_mapping) with a single Tool dataclass and a TOOLS list, so each tool's schema and callable live under one name and can't drift apart. - Add Tool(name, description, parameters, run) with a .schema() method; define SEARCH_TOOL and ASK_DATABASE_TOOL instances and a single TOOLS list as the source of truth. Adding a tool is appending one Tool. - Build the ask_database schema from Medication._meta (concrete_fields / db_table) instead of introspecting the live database, removing the import-time DB query and the deferred database_schema_string import. - Bind the request user at dispatch time: invoke_functions_from_response and handle_tool_calls_with_reasoning now take (tools, user), index tools by name, and call tool.run(user=user, **arguments). This drops make_tool_mapping / make_search_tool_mapping and their closures. - assistant_services builds the schema list with [tool.schema() for tool in TOOLS] and forwards TOOLS + user to the loop. - Update tests to cover the Tool instances and pass (tools, user) to the loop; the tool_services.search_documents / ask_database patch paths still resolve.
Convert every relative import in api/views/assistant to its absolute equivalent (assistant_services, search_tool, tool_services, urls, views). Multi-dot forms like `...services.tools.database` and `..listMeds.models` were error-prone to read and would break silently if a file's package depth changed; the absolute paths are move-safe and unambiguous. Expand the tool_services.py import comments to record why search_documents and ask_database are imported as bare names: the tests patch them at their use site (api.views.assistant.tool_services.<name>), not their definition site, so mock.patch rebinds the reference SEARCH_TOOL.run actually resolves at call time. Note the maintenance guard — qualifying those calls would move the patch target and break the tests. Move the _medication_schema_string helper to sit directly above its only caller, ASK_DATABASE_TOOL, instead of above SEARCH_TOOL. No behavior change: all bound names and patch targets are preserved.
Replace the _medication_schema_string() helper (which read columns from Medication._meta) with a hand-written _MEDICATION_SCHEMA_STRING constant, and drop the now-unused Medication import. The _meta approach auto-synced with the model but dumped every column and pulled in an app-registry dependency (AppRegistryNotReady if imported during app startup). For a 4-column, stable table, a curated constant is simpler, lets us hide columns from the LLM (omit `id`, which it never filters on), and removes the startup coupling — at the cost of a one-line manual update if the table's columns ever change, which the comment calls out. Note: this drops `id` from the schema the model sees (intentional curation, not just a port of the old behavior). Also document in the Tool docstring why behavior is a `run` field (composition) rather than a subclass method: the tools differ only in which function runs, so they are instances of one concept, not distinct types. Add a TODO listing the signals that would justify flipping to Tool(ABC) + per-tool subclasses (per-type state, overriding more than run, or a per-type/abstractmethod-enforced contract).
The eval showed what the assistant said but not how it chose tools. Tool-call info was produced in invoke_functions_from_response and dropped at every return boundary; caught tool exceptions were fed back to the model as strings, so a run where ask_database threw every question read as clean rows. Carry it up as return values (over a mutable out-param: honest domain data that grows via defaulted fields, no call-site churn): - agentic_loop: add ToolCallStatus (OK/FAILED/UNREGISTERED — a bool was both redundant with error and lossy), ToolCall(name, status, arguments, output, error), and AssistantResult(output_text, response_id, tool_calls). invoke_functions_from_response returns (messages, tool_calls); the loop accumulates across iterations and returns AssistantResult. - assistant_services: return AssistantResult (pass-through). Drops the TODO. - views: read result fields; JSON body unchanged. - eval_assistant: run_one times the call locally and adds tools_called, tool_call_count, tool_error_count, tool_calls_json, response_id, duration_s — making tool_error_count > 0 while error is None visible. Deferred: token-cost, turn count, correctness scoring. Tests updated.
Hoist MODEL_NAME into assistant_services, import it in eval_assistant, and drop the duplicated literal — the only functional change here. The rest is comments. Token usage and turn count, the scoring layer, and the INSTRUCTIONS sidecar were each designed in this pass and deliberately not built; the TODOs sit where the work will happen and carry the reasoning.
The eval had never been run. Every test mocks run_assistant, so a green suite proved run_one's row shaping but never that a CSV came out. Three defects stopped it running at all: the sys.path depth, pandas missing from the backend image, and the uv shebang and PEP 723 header, which never described a runnable configuration. A fourth is different in kind, and only the run could surface it. The cold-start race on the embedding model does not stop the eval — it corrupts it, completing normally and writing a CSV that looks clean. Warming the model before the pool addresses it here.
…run-unblocking fixes
Removed. None passes the mutation test — can a plausible one-line production
change turn it red *and* ship a bug?
- test_ask_database_tool_run_ignores_user: single-arg forward; the wrong
version raises TypeError on first call.
- test_tools_registry_contains_both_tools: restated the TOOLS literal, so
appending a tool was also how you broke the test.
- test_run_assistant_sends_message_as_user_input: echoed a hardcoded dict.
- test_run_assistant_forwards_tools_and_user_to_loop: `args[3] is TOOLS`,
coupled to argument position.
Merged three groups into parametrize tables. Two gain coverage:
- The loop test now checks the previous_response_id chain per turn, not just
the first follow-up.
- The FAILED/UNREGISTERED table exposes that `arguments` is parsed only
inside the registered branch.
Tests whose assertions differ in kind stayed separate.
Added — flagged because they ride along with a commit that is otherwise
subtraction:
- test_failed_status_... dispatched a MagicMock named "search_documents", so
it duplicated the error table. It now dispatches the real SEARCH_TOOL, and
is the only test proving the warm-up and search_tool.py fixes meet.
- test_run_one_row_carries_every_csv_column: DictWriter raises on an extra
key but fills a missing one with restval (""), so a column forgotten in one
run_one row literal reaches the CSV as an empty cell, not an error.
Also reworded a tool_services.py comment asserting that tests patch
ask_database there — true until this commit.
Accepted: user -> run_assistant -> loop is no longer asserted. Bare positional
forward, no decision in it, both adjacent legs still covered.
22 tests -> 15 functions / 19 cases. Per-test rationale is in the docstrings
The new names say waht the functions do rather than borrowing the OpenAI Cookbook vocab
Errors in the run or values for token usage and tool calls are to be None, not 0
"Turn" now means one whole run_assistant call and "iteration" means one pass of the agentic loop (one responses.create). Removed: model, response_id, tools_called, turn_count
Dropped tests that cover telemetry, or failures that are loud on the first real call Each case was checked by deliberately breaking the code using Claude Opus 5.5
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The public SQL tool permits allowlist bypasses and reports database failures as successful calls.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (6)
Public SQL endpoint allows unauthorized table access · New System instructions do not permit database tool results · New CSV schema omits response ID and model attribution · New Deleted tests leave database and evaluation behavior untested · New Database errors are swallowed and reported as successful · New Fix typo in explanatory comment · New
What changed in this PR
Adds database-backed medication research and tool-call telemetry to the assistant.
Changes:
- Introduces a unified tool registry and
ask_database. - Extracts the agentic loop and document search.
- Expands evaluation telemetry and consolidates tests.
| File | Description |
|---|---|
views.py |
Adapts API response to AgentResult. |
urls.py |
Uses an absolute view import. |
tool_services.py |
Registers search and database tools. |
test_tool_services.py |
Removes prior tool tests. |
test_eval_assistant.py |
Removes prior evaluation tests. |
test_assistant_services.py |
Removes prior service tests. |
test_agentic_loop.py |
Adds consolidated loop tests. |
search_tool.py |
Extracts semantic document search. |
eval_assistant.py |
Adds CSV telemetry and concurrent evaluation. |
assistant_types.py |
Defines tool, result, and telemetry types. |
assistant_services.py |
Integrates the unified tools and loop. |
assistant_prompts.py |
Documents unresolved prompt/tool conflict. |
agentic_loop.py |
Implements dispatch, statuses, and usage tracking. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| (e.g. LOWER(name) = LOWER('lurasidone')). The input must be a single, fully-formed | ||
| SQL SELECT query. |
There was a problem hiding this comment.
Agreed, the substring checks can be bypassed. This PR now unregisters ask_database from the assistant, and making it safe (single statement, real table allowlist, read-only DB role) is high-priority future work
| # TODO: Mention ask_database — the prompt names only search_documents and says to "ALWAYS use" | ||
| # it first, steering the model away from ask_database |
There was a problem hiding this comment.
With ask_database unregistered, there's no conflict in this PR. Updating the prompt and tool description is future work, so its effect on answers can be measured on its own with the eval.
| FIELDNAMES = [ | ||
| "branch", | ||
| "question", | ||
| "error", | ||
| "response_output_text", | ||
| "duration_s", | ||
| # No total_tokens column: it is total_input_tokens + total_output_tokens | ||
| "total_input_tokens", | ||
| "total_cached_input_tokens", | ||
| "total_output_tokens", | ||
| "total_reasoning_output_tokens", | ||
| "total_tool_calls", | ||
| "total_tool_errors", | ||
| "tool_calls_json", | ||
| "token_usages_json", |
There was a problem hiding this comment.
Dropping both columns was deliberate, to keep the CSV to what current analysis uses. I've corrected the MODEL_NAME comment that said otherwise, and adding them back is future work.
| # Deliberately not tested here (lower stakes, or fails loudly on the first real call): | ||
| # eval and token telemetry, previous_response_id handling in run_assistant, tool | ||
| # schemas and adapters, and the prompt text (model behavior, measured by the eval). |
There was a problem hiding this comment.
This PR deliberately replaced the old suite with 3 tests covering the riskiest loop behavior (grounding answers in tool output, reporting tool failures, and scoping retrieval to the request user). The description now says 3, not 19, and eval tests are future work
| # ask_database queries the shared medication table, so it ignores the request user. | ||
| run=lambda user, query: ask_database(query), |
There was a problem hiding this comment.
Confirmed. With ask_database unregistered it can't affect the assistant in this PR, and making it raise on failure, so failures record as FAILED, is listed as high-priority future work before it's registered again
|
|
||
| # The schema string describing the queryable medication table for ask_database's prompt. | ||
| # Kept in sync by hand with api.views.listMeds.models.Medication rather than deriving it from | ||
| # Django's Model._meta becuase the table is small and stable |
ask_database's guards only check that a query starts with "select" and contains "from api_medication", so a UNION or a second statement can reach other tables, and the assistant endpoint is AllowAny It also returns errors as text, so a failed query would be recorded as OK



Description
This PR gets the assistant's tool loop working, records every tool call and how many tokens it used, and adds an eval script so we can measure answers. Search failures now show up as FAILED instead of turning into confident answers. The endpoint's request and response are unchanged.
Related Issue
#521
Manual Tests
docker compose exec -e EVAL_BRANCH=521-research-agent-tools backend python api/views/assistant/eval_assistant.py: 5 questions, 0 errors, and the token totals match the per-call records. Paste the run summary from the re-run on the final commit; see checklist item 2.
Log:
CSV: 521-research-agent-tools-20261005T215441.csv
Results:
search_documentscall that succeeded, andask_databasenever came up.ask_databaseshows up in the numbers. The first call per question now sends about 125 fewer input tokens (≈660 → ≈535), roughly the size of the tool definition we stopped sending.Logged Out:

Logged In:
Automated Tests
docker compose exec backend pytest api/views/assistant -v: 3 passed. The old suite was replaced with 3 tests covering the highest-risk behaviour: tool output reaching the model, tool failures being reported, and search being scoped to the request user. Each was checked by deliberately breaking the code.
Documentation
There are no user-facing doc changes, and the API schema is unchanged. Code comments and TODOs mark each future-work item, and the eval script's header documents how to run it.
Reviewers
Notes
ask_database is defined but deliberately not registered, because its SQL checks can be bypassed (see Copilot's review). Paste the Future work section from WORKLOG.md below this line.