Skip to content

[#521] Research Agent Tools - #524

Merged
sahilds1 merged 26 commits into
CodeForPhilly:developfrom
sahilds1:521-research-agent-tools
Oct 5, 2026
Merged

sahilds1 merged 26 commits into
CodeForPhilly:developfrom
sahilds1:521-research-agent-tools

Conversation

@sahilds1

@sahilds1 sahilds1 commented Jul 15, 2026 •

Copy link
Copy Markdown
Collaborator

Description

This PR gets the assistant's tool loop working, records every tool call and how many tokens it used, and adds an eval script so we can measure answers. Search failures now show up as FAILED instead of turning into confident answers. The endpoint's request and response are unchanged.

Related Issue

#521

Manual Tests

docker compose exec -e EVAL_BRANCH=521-research-agent-tools backend python api/views/assistant/eval_assistant.py: 5 questions, 0 errors, and the token totals match the per-call records. Paste the run summary from the re-run on the final commit; see checklist item 2.

Log:

➜  balancer-main git:(521-research-agent-tools) ✗ docker compose exec -e EVAL_BRANCH=521-research-agent-tools backend python api/views/assistant/eval_assistant.py
2026-10-05 21:54:22,582 - INFO - Starting evaluation: branch=521-research-agent-tools, model=gpt-5-nano, questions=5
2026-10-05 21:54:22,582 - INFO - Loading SentenceTransformer model
2026-10-05 21:54:22,582 - INFO - Use pytorch device_name: cpu
2026-10-05 21:54:22,582 - INFO - Load pretrained SentenceTransformer: paraphrase-MiniLM-L6-v2
2026-10-05 21:54:22,658 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/modules.json "HTTP/1.1 307 Temporary Redirect"
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2026-10-05 21:54:22,658 - WARNING - Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2026-10-05 21:54:22,678 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/modules.json "HTTP/1.1 200 OK"
2026-10-05 21:54:22,712 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/config_sentence_transformers.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,727 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/config_sentence_transformers.json "HTTP/1.1 200 OK"
2026-10-05 21:54:22,759 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/config_sentence_transformers.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,776 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/config_sentence_transformers.json "HTTP/1.1 200 OK"
2026-10-05 21:54:22,806 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/README.md "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,822 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/README.md "HTTP/1.1 200 OK"
2026-10-05 21:54:22,852 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/modules.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,868 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/modules.json "HTTP/1.1 200 OK"
2026-10-05 21:54:22,896 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/sentence_bert_config.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,910 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/sentence_bert_config.json "HTTP/1.1 200 OK"
2026-10-05 21:54:22,946 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/adapter_config.json "HTTP/1.1 404 Not Found"
2026-10-05 21:54:22,977 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:22,996 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/config.json "HTTP/1.1 200 OK"
Loading weights: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 103/103 [00:00<00:00, 8762.95it/s]
BertModel LOAD REPORT from: sentence-transformers/paraphrase-MiniLM-L6-v2
Key                     | Status     |  |
------------------------+------------+--+-
embeddings.position_ids | UNEXPECTED |  |

Notes:
- UNEXPECTED	:can be ignored when loading from different task/architecture; not ok if you expect identical arch.
2026-10-05 21:54:23,067 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:23,086 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/config.json "HTTP/1.1 200 OK"
2026-10-05 21:54:23,116 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:23,131 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/tokenizer_config.json "HTTP/1.1 200 OK"
2026-10-05 21:54:23,163 - INFO - HTTP Request: GET https://huggingface.co/api/models/sentence-transformers/paraphrase-MiniLM-L6-v2/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found"
2026-10-05 21:54:23,202 - INFO - HTTP Request: GET https://huggingface.co/api/models/sentence-transformers/paraphrase-MiniLM-L6-v2/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK"
2026-10-05 21:54:23,275 - INFO - HTTP Request: HEAD https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2/resolve/main/1_Pooling/config.json "HTTP/1.1 307 Temporary Redirect"
2026-10-05 21:54:23,291 - INFO - HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/sentence-transformers/paraphrase-MiniLM-L6-v2/c9a2bfebc254878aee8c3aca9e6844d5bbb102d1/1_Pooling%2Fconfig.json "HTTP/1.1 200 OK"
2026-10-05 21:54:23,328 - INFO - HTTP Request: GET https://huggingface.co/api/models/sentence-transformers/paraphrase-MiniLM-L6-v2 "HTTP/1.1 200 OK"
2026-10-05 21:54:26,174 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:26,267 - INFO - Invoking tool: search_documents with arguments: {'query': 'lithium kidney disease risks renal impairment patients with kidney disease lithium toxicity CKD risk monitoring guidelines'}
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 13.60it/s]
2026-10-05 21:54:26,364 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:26,389 - INFO - Invoking tool: search_documents with arguments: {'query': 'medications recommended for bipolar depression'}
Batches:   0%|                                                                                                                                                   | 0/1 [00:00<?, ?it/s]2026-10-05 21:54:26,393 - INFO - Tool search_documents completed successfully
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 54.43it/s]
2026-10-05 21:54:26,437 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:26,453 - INFO - Tool search_documents completed successfully
2026-10-05 21:54:26,480 - INFO - Invoking tool: search_documents with arguments: {'query': 'valproate versus lithium mania comparison randomized controlled trial efficacy lithium valproate mania bipolar disorder information source'}
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 54.27it/s]
2026-10-05 21:54:26,528 - INFO - Tool search_documents completed successfully
2026-10-05 21:54:26,965 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:27,005 - INFO - Invoking tool: search_documents with arguments: {'query': 'quetiapine bipolar disorder evidence efficacy safety randomized trials guidelines meta-analysis queue quetiapine XR bipolar'}
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 37.33it/s]
2026-10-05 21:54:27,068 - INFO - Tool search_documents completed successfully
2026-10-05 21:54:28,170 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:28,184 - INFO - Invoking tool: search_documents with arguments: {'query': 'mood stabilizers safe during pregnancy pregnancy safety lamotrigine valproate lithium lithium during pregnancy pregnancy safety bipolar disorder sources'}
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 30.05it/s]
2026-10-05 21:54:28,251 - INFO - Tool search_documents completed successfully
2026-10-05 21:54:32,944 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:35,217 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:38,267 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:38,865 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:41,831 - INFO - HTTP Request: POST https://api.openai.com/v1/responses "HTTP/1.1 200 OK"
2026-10-05 21:54:41,905 - INFO - Results saved to /usr/src/server/api/views/assistant/results/521-research-agent-tools-20261005T215441.csv

CSV: 521-research-agent-tools-20261005T215441.csv

Results:

  • All 5 questions ran with no errors. Each one made a single search_documents call that succeeded, and ask_database never came up.
  • Token numbers add up. Every per-question total matches its per-call records, and cached and reasoning tokens always stay within their totals.
  • Unregistering ask_database shows up in the numbers. The first call per question now sends about 125 fewer input tokens (≈660 → ≈535), roughly the size of the tool definition we stopped sending.
  • No cached tokens this run, which is expected. Each question needed only two OpenAI calls, and caching only starts once a repeated prompt is long enough.
  • The lithium and kidney disease answer is honest. It says it found nothing, and the search results really didn't mention the kidney.
  • Each question took 10–19 seconds, and the whole run about 19 seconds with 5 running in parallel.

Logged Out:
Screenshot 2026-10-05 at 6 20 35 PM

Screenshot 2026-10-05 at 6 24 57 PM

Logged In:

Screenshot 2026-10-05 at 6 29 27 PM Screenshot 2026-10-05 at 6 30 56 PM

Automated Tests

docker compose exec backend pytest api/views/assistant -v: 3 passed. The old suite was replaced with 3 tests covering the highest-risk behaviour: tool output reaching the model, tool failures being reported, and search being scoped to the request user. Each was checked by deliberately breaking the code.

➜  balancer-main git:(521-research-agent-tools) ✗ docker compose exec backend pytest -v
================================================================================= test session starts =================================================================================
platform linux -- Python 3.11.4, pytest-9.0.2, pluggy-1.6.0 -- /usr/local/bin/python
cachedir: .pytest_cache
django: version: 4.2.3, settings: balancer_backend.settings (from ini)
rootdir: /usr/src/server
configfile: pytest.ini
plugins: django-4.12.0, anyio-4.12.1
collected 32 items

api/services/test_embedding_services.py::test_build_query_authenticated_uses_or_filter PASSED                                                                                   [  3%]
api/services/test_embedding_services.py::test_build_query_unauthenticated_uses_superuser_only_filter PASSED                                                                     [  6%]
api/services/test_embedding_services.py::test_build_query_annotates_and_orders_by_distance PASSED                                                                               [  9%]
api/services/test_embedding_services.py::test_build_query_no_document_filter_when_both_none PASSED                                                                              [ 12%]
api/services/test_embedding_services.py::test_build_query_guid_takes_precedence_over_document_name PASSED                                                                       [ 15%]
api/services/test_embedding_services.py::test_build_query_guid_filter_applied PASSED                                                                                            [ 18%]
api/services/test_embedding_services.py::test_build_query_document_name_filter_applied PASSED                                                                                   [ 21%]
api/services/test_embedding_services.py::test_build_query_empty_string_guid_falls_back_to_document_name PASSED                                                                  [ 25%]
api/services/test_embedding_services.py::test_build_query_respects_num_results PASSED                                                                                           [ 28%]
api/services/test_embedding_services.py::test_build_query_returns_unevaluated_queryset PASSED                                                                                   [ 31%]
api/services/test_embedding_services.py::test_evaluate_query_empty_queryset PASSED                                                                                              [ 34%]
api/services/test_embedding_services.py::test_evaluate_query_maps_fields PASSED                                                                                                 [ 37%]
api/services/test_embedding_services.py::test_evaluate_query_none_upload_file PASSED                                                                                            [ 40%]
api/services/test_embedding_services.py::test_log_usage_empty_results PASSED                                                                                                    [ 43%]
api/services/test_embedding_services.py::test_log_usage_unauthenticated_user_stored_as_none PASSED                                                                              [ 46%]
api/services/test_embedding_services.py::test_log_usage_none_user_stored_as_none PASSED                                                                                         [ 50%]
api/services/test_embedding_services.py::test_log_usage_computes_distance_stats PASSED                                                                                          [ 53%]
api/services/test_embedding_services.py::test_log_usage_swallows_exceptions PASSED                                                                                              [ 56%]
api/services/test_embedding_services.py::test_get_closest_embeddings_wiring PASSED                                                                                              [ 59%]
api/views/assistant/test_agentic_loop.py::test_loop_feeds_each_tool_output_back_until_the_model_answers PASSED                                                                  [ 62%]
api/views/assistant/test_agentic_loop.py::test_tool_failures_reach_the_model_and_are_recorded_by_kind PASSED                                                                    [ 65%]
api/views/assistant/test_agentic_loop.py::test_search_uses_the_request_user_and_tells_failure_from_no_match PASSED                                                              [ 68%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_falls_back_to_chatgpt_if_no_title_found PASSED                                                                      [ 71%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_falls_back_to_font_size_if_metadata_title_does_not_match_regex PASSED                                               [ 75%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_falls_back_to_font_size_if_metadata_title_is_empty PASSED                                                           [ 78%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_font_size_finds_title_on_later_page PASSED                                                                          [ 81%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_font_size_ignores_short_spans PASSED                                                                                [ 84%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_font_size_joins_adjacent_spans_in_same_block PASSED                                                                 [ 87%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_font_size_returns_none_when_no_regex_match PASSED                                                                   [ 90%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_prefers_metadata_title_if_valid PASSED                                                                              [ 93%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_strips_quotes_from_openai_title PASSED                                                                              [ 96%]
api/views/uploadFile/test_title.py::TestGenerateTitle::test_truncates_long_openai_title PASSED                                                                                  [100%]

================================================================================= 32 passed in 1.87s ==================================================================================

Documentation

There are no user-facing doc changes, and the API schema is unchanged. Code comments and TODOs mark each future-work item, and the eval script's header documents how to run it.

Reviewers

Notes

ask_database is defined but deliberately not registered, because its SQL checks can be bypassed (see Copilot's review). Paste the Future work section from WORKLOG.md below this line.

@sahilds1 sahilds1 self-assigned this Jul 15, 2026
@sahilds1 sahilds1 changed the title Research Agent Tools [WIP] [#521] Research Agent Tools Jul 15, 2026
@sahilds1
sahilds1 marked this pull request as draft July 15, 2026 22:50
sahilds1 added 23 commits July 16, 2026 15:10
Add ask_database as a second tool for the assistant, alongside the
existing semantic search_documents tool.

- Reuse the SELECT-only, ALLOWED_TABLES-guarded ask_database
  implementation from services/tools/database.py rather than
  reimplementing the query guards in the assistant.
- Add get_tools_schema() / make_tool_mapping() as an aggregation seam
  in tool_services.py so assistant_services.py no longer names
  individual tools; new tools are registered in one place.
- Build the ask_database schema in the flattened Responses-API shape
  (not the nested Chat Completions shape from services/tools), and
  defer the database_schema_string import to call time so importing
  the module never triggers a DB query.
- Split tool_services.py: move search_documents into search_tool.py and
  the agentic-loop helpers (handle_tool_calls_with_reasoning,
  invoke_functions_from_response) into agentic_loop.py, adding the
  imports each module needs.
- Point importers (assistant_services.py, test_tool_services.py) at the
  defining module for each symbol instead of re-exporting through
  tool_services.py.
- Document the search/SQL tool overlap risk and the import-time DB
  access caveat inline.
Replace the two parallel tool registries (get_tools_schema /
make_tool_mapping) with a single Tool dataclass and a TOOLS list, so each
tool's schema and callable live under one name and can't drift apart.

- Add Tool(name, description, parameters, run) with a .schema() method;
  define SEARCH_TOOL and ASK_DATABASE_TOOL instances and a single TOOLS
  list as the source of truth. Adding a tool is appending one Tool.
- Build the ask_database schema from Medication._meta (concrete_fields /
  db_table) instead of introspecting the live database, removing the
  import-time DB query and the deferred database_schema_string import.
- Bind the request user at dispatch time: invoke_functions_from_response
  and handle_tool_calls_with_reasoning now take (tools, user), index
  tools by name, and call tool.run(user=user, **arguments). This drops
  make_tool_mapping / make_search_tool_mapping and their closures.
- assistant_services builds the schema list with
  [tool.schema() for tool in TOOLS] and forwards TOOLS + user to the loop.
- Update tests to cover the Tool instances and pass (tools, user) to the
  loop; the tool_services.search_documents / ask_database patch paths
  still resolve.
Convert every relative import in api/views/assistant to its absolute
equivalent (assistant_services, search_tool, tool_services, urls, views).
Multi-dot forms like `...services.tools.database` and `..listMeds.models`
were error-prone to read and would break silently if a file's package
depth changed; the absolute paths are move-safe and unambiguous.

Expand the tool_services.py import comments to record why search_documents
and ask_database are imported as bare names: the tests patch them at their
use site (api.views.assistant.tool_services.<name>), not their definition
site, so mock.patch rebinds the reference SEARCH_TOOL.run actually resolves
at call time. Note the maintenance guard — qualifying those calls would
move the patch target and break the tests.

Move the _medication_schema_string helper to sit directly above its only
caller, ASK_DATABASE_TOOL, instead of above SEARCH_TOOL.

No behavior change: all bound names and patch targets are preserved.
Replace the _medication_schema_string() helper (which read columns from
Medication._meta) with a hand-written _MEDICATION_SCHEMA_STRING constant,
and drop the now-unused Medication import. The _meta approach auto-synced
with the model but dumped every column and pulled in an app-registry
dependency (AppRegistryNotReady if imported during app startup). For a
4-column, stable table, a curated constant is simpler, lets us hide
columns from the LLM (omit `id`, which it never filters on), and removes
the startup coupling — at the cost of a one-line manual update if the
table's columns ever change, which the comment calls out.

Note: this drops `id` from the schema the model sees (intentional
curation, not just a port of the old behavior).

Also document in the Tool docstring why behavior is a `run` field
(composition) rather than a subclass method: the tools differ only in
which function runs, so they are instances of one concept, not distinct
types. Add a TODO listing the signals that would justify flipping to
Tool(ABC) + per-tool subclasses (per-type state, overriding more than
run, or a per-type/abstractmethod-enforced contract).
The eval showed what the assistant said but not how it chose tools. Tool-call
info was produced in invoke_functions_from_response and dropped at every return
boundary; caught tool exceptions were fed back to the model as strings, so a run
where ask_database threw every question read as clean rows.

Carry it up as return values (over a mutable out-param: honest domain data that
grows via defaulted fields, no call-site churn):

- agentic_loop: add ToolCallStatus (OK/FAILED/UNREGISTERED — a bool was both
  redundant with error and lossy), ToolCall(name, status, arguments, output,
  error), and AssistantResult(output_text, response_id, tool_calls).
  invoke_functions_from_response returns (messages, tool_calls); the loop
  accumulates across iterations and returns AssistantResult.
- assistant_services: return AssistantResult (pass-through). Drops the TODO.
- views: read result fields; JSON body unchanged.
- eval_assistant: run_one times the call locally and adds tools_called,
  tool_call_count, tool_error_count, tool_calls_json, response_id, duration_s —
  making tool_error_count > 0 while error is None visible.

Deferred: token-cost, turn count, correctness scoring. Tests updated.
Hoist MODEL_NAME into assistant_services, import it in eval_assistant, and drop
the duplicated literal — the only functional change here.

The rest is comments. Token usage and turn count, the scoring layer, and the
INSTRUCTIONS sidecar were each designed in this pass and deliberately not built;
the TODOs sit where the work will happen and carry the reasoning.
The eval had never been run. Every test mocks run_assistant, so a green
suite proved run_one's row shaping but never that a CSV came out.

Three defects stopped it running at all: the sys.path depth, pandas
missing from the backend image, and the uv shebang and PEP 723 header,
which never described a runnable configuration.

A fourth is different in kind, and only the run could surface it. The
cold-start race on the embedding model does not stop the eval — it
corrupts it, completing normally and writing a CSV that looks clean.
Warming the model before the pool addresses it here.
Removed. None passes the mutation test — can a plausible one-line production
change turn it red *and* ship a bug?
  - test_ask_database_tool_run_ignores_user: single-arg forward; the wrong
    version raises TypeError on first call.
  - test_tools_registry_contains_both_tools: restated the TOOLS literal, so
    appending a tool was also how you broke the test.
  - test_run_assistant_sends_message_as_user_input: echoed a hardcoded dict.
  - test_run_assistant_forwards_tools_and_user_to_loop: `args[3] is TOOLS`,
    coupled to argument position.

Merged three groups into parametrize tables. Two gain coverage:
  - The loop test now checks the previous_response_id chain per turn, not just
    the first follow-up.
  - The FAILED/UNREGISTERED table exposes that `arguments` is parsed only
    inside the registered branch.
Tests whose assertions differ in kind stayed separate.

Added — flagged because they ride along with a commit that is otherwise
subtraction:
  - test_failed_status_... dispatched a MagicMock named "search_documents", so
    it duplicated the error table. It now dispatches the real SEARCH_TOOL, and
    is the only test proving the warm-up and search_tool.py fixes meet.
  - test_run_one_row_carries_every_csv_column: DictWriter raises on an extra
    key but fills a missing one with restval (""), so a column forgotten in one
    run_one row literal reaches the CSV as an empty cell, not an error.

Also reworded a tool_services.py comment asserting that tests patch
ask_database there — true until this commit.

Accepted: user -> run_assistant -> loop is no longer asserted. Bare positional
forward, no decision in it, both adjacent legs still covered.

22 tests -> 15 functions / 19 cases. Per-test rationale is in the docstrings
The new names say waht the functions do rather than borrowing the OpenAI Cookbook vocab
Errors in the run or values for  token usage and tool calls are to be None, not 0
"Turn" now means one whole run_assistant call and "iteration" means one
pass of the agentic loop (one responses.create).

Removed: model, response_id, tools_called, turn_count
Dropped tests that  cover telemetry, or failures that are loud on the first real call

Each case was checked by deliberately breaking the code using Claude Opus 5.5

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The public SQL tool permits allowlist bypasses and reports database failures as successful calls.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 4 Medium severity · 1 Low severity

Open (6)
What changed in this PR

Adds database-backed medication research and tool-call telemetry to the assistant.

Changes:

  • Introduces a unified tool registry and ask_database.
  • Extracts the agentic loop and document search.
  • Expands evaluation telemetry and consolidates tests.
File Description
views.py Adapts API response to AgentResult.
urls.py Uses an absolute view import.
tool_services.py Registers search and database tools.
test_tool_services.py Removes prior tool tests.
test_eval_assistant.py Removes prior evaluation tests.
test_assistant_services.py Removes prior service tests.
test_agentic_loop.py Adds consolidated loop tests.
search_tool.py Extracts semantic document search.
eval_assistant.py Adds CSV telemetry and concurrent evaluation.
assistant_types.py Defines tool, result, and telemetry types.
assistant_services.py Integrates the unified tools and loop.
assistant_prompts.py Documents unresolved prompt/tool conflict.
agentic_loop.py Implements dispatch, statuses, and usage tracking.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +47 to +48
(e.g. LOWER(name) = LOWER('lurasidone')). The input must be a single, fully-formed
SQL SELECT query.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, the substring checks can be bypassed. This PR now unregisters ask_database from the assistant, and making it safe (single statement, real table allowlist, read-only DB role) is high-priority future work

Comment on lines +11 to +12
# TODO: Mention ask_database — the prompt names only search_documents and says to "ALWAYS use"
# it first, steering the model away from ask_database

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With ask_database unregistered, there's no conflict in this PR. Updating the prompt and tool description is future work, so its effect on answers can be measured on its own with the eval.

Comment on lines +38 to +52
FIELDNAMES = [
"branch",
"question",
"error",
"response_output_text",
"duration_s",
# No total_tokens column: it is total_input_tokens + total_output_tokens
"total_input_tokens",
"total_cached_input_tokens",
"total_output_tokens",
"total_reasoning_output_tokens",
"total_tool_calls",
"total_tool_errors",
"tool_calls_json",
"token_usages_json",

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropping both columns was deliberate, to keep the CSV to what current analysis uses. I've corrected the MODEL_NAME comment that said otherwise, and adding them back is future work.

Comment on lines +4 to +6
# Deliberately not tested here (lower stakes, or fails loudly on the first real call):
# eval and token telemetry, previous_response_id handling in run_assistant, tool
# schemas and adapters, and the prompt text (model behavior, measured by the eval).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR deliberately replaced the old suite with 3 tests covering the riskiest loop behavior (grounding answers in tool output, reporting tool failures, and scoping retrieval to the request user). The description now says 3, not 19, and eval tests are future work

Comment on lines +70 to +71
# ask_database queries the shared medication table, so it ignores the request user.
run=lambda user, query: ask_database(query),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed. With ask_database unregistered it can't affect the assistant in this PR, and making it raise on failure, so failures record as FAILED, is listed as high-priority future work before it's registered again


# The schema string describing the queryable medication table for ask_database's prompt.
# Kept in sync by hand with api.views.listMeds.models.Medication rather than deriving it from
# Django's Model._meta becuase the table is small and stable

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, thanks.

ask_database's guards only check that a query starts with "select"
and contains "from api_medication", so a UNION or a second statement
can reach other tables, and the assistant endpoint is AllowAny

It also returns errors as text, so a failed query would be recorded as OK
@sahilds1
sahilds1 marked this pull request as ready for review October 5, 2026 20:24
@sahilds1 sahilds1 changed the title [WIP] [#521] Research Agent Tools [#521] Research Agent Tools Oct 5, 2026
@sahilds1 sahilds1 mentioned this pull request Oct 5, 2026
3 tasks
@sahilds1
sahilds1 merged commit 3bd48f4 into CodeForPhilly:develop Oct 5, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants