Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
d897194
Fix mobile nav 'Medication Suggester' link target and label
c-tonneslan May 28, 2026
cdd6ba4
docs: add local frontend and API notes (localhost:3000/8000)
snaeem3 Jun 25, 2026
9f71fbd
feat: update DB config defaults from RDS to CloudNativePG
TineoC Jun 27, 2026
a1ab4ee
fix: resolve Ruff lint errors blocking CI
TineoC Jul 10, 2026
3171cc1
Merge pull request #515 from snaeem3/update/readme-local-servers
sahilds1 Jul 15, 2026
b3669b0
Add TODOs for adding tools
sahilds1 Jul 15, 2026
b5710e9
Add ask_database tool and split tool module
sahilds1 Jul 16, 2026
aba10f8
Refactor model tools as a Tool dataclass, drop the aggregators
sahilds1 Jul 17, 2026
76a78b0
Use absolute imports in assistant module; clarify test-patching contract
sahilds1 Jul 21, 2026
35a360d
Hardcode ask_database schema string; document Tool composition choice
sahilds1 Jul 21, 2026
04e7b2d
Merge pull request #512 from c-tonneslan/fix/494-mobile-nav-medicatio…
sahilds1 Jul 21, 2026
100ed12
Capture tool calls and duration in the eval CSV
sahilds1 Jul 24, 2026
85a1a0a
Share one MODEL_NAME constant; record deferred eval work as TODOs
sahilds1 Jul 27, 2026
dceec41
Add 'No' value to radio button
snaeem3 Jul 27, 2026
179f724
Merge pull request #514 from CodeForPhilly/feature/migrate-rds-to-cnp…
TineoC Jul 28, 2026
7de84ab
Merge pull request #525 from snaeem3/388-suicide-risk-button
sahilds1 Aug 4, 2026
53afce8
Merge branch 'develop' into 521-research-agent-tools
sahilds1 Aug 4, 2026
b6eca89
Unblock the eval's first end-to-end run
sahilds1 Aug 4, 2026
2e58e11
Eval reported clean runs while retreivals crashed because of a race
sahilds1 Aug 4, 2026
f129ce4
Note: b6eca89 bundles the race on the embedding model fixes with the …
sahilds1 Aug 4, 2026
ff75d34
Test only the logic we wrote: drop glue tests, dedupe with parametrize
sahilds1 Aug 7, 2026
408b2a4
Add comments for open follow ups
sahilds1 Aug 11, 2026
a4949cb
Tool, ToolCallStatus, ToolCall and AssistantResult now live in one mo…
sahilds1 Aug 11, 2026
fba44f2
invoke_functions_from_response's function_call branch moves to _execu…
sahilds1 Aug 11, 2026
72582ab
Note: fba44f2 bundles renaming initial_response fix
sahilds1 Aug 11, 2026
fb4ab35
Rename the agentic loop's functions and record types
sahilds1 Aug 16, 2026
daed3de
DOC: Condense comments in assistant
sahilds1 Aug 17, 2026
76f5857
Placeholder TODOs for token capture
sahilds1 Sep 1, 2026
b6c4c4d
DOC TODOs before implementing token usage capture
sahilds1 Sep 15, 2026
5050183
Placeholder implementation for token capture
sahilds1 Sep 16, 2026
d6c468c
DOC Comments for placeholder implementation of token usage
sahilds1 Sep 16, 2026
46d0b41
Rename TurnUsage to TokenUsage and reshape the eval CSV
sahilds1 Sep 25, 2026
0b13504
Replace the assistant test suite with high priority tests
sahilds1 Sep 25, 2026
8c5b686
Unregister ask_database from the assistant
sahilds1 Oct 5, 2026
e3d0f0c
Add TODOs for remaining work and how per-iteration data is recorded
sahilds1 Oct 5, 2026
3bd48f4
Merge pull request #524 from sahilds1/521-research-agent-tools
sahilds1 Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,12 @@ Tools used for development:

Start the Postgres, Django REST, and React services by starting Docker Desktop and running `docker compose up --build`

Local development servers:

- **Frontend (React dev server):** http://localhost:3000 — when developing locally, open the React app at this address.
- **Backend / API (Django):** http://localhost:8000
- *Note:* if you open http://localhost:8000 in a browser you will see a minimal single-page fallback for the frontend, which is non-functional for development. To view the working frontend during development, use http://localhost:3000.

#### Postgres

The application supports connecting to PostgreSQL databases via:
Expand Down
9 changes: 5 additions & 4 deletions config/env/dev.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,14 @@ SQL_PASSWORD=balancer

# Connection Type Examples:
#
# CloudNativePG (Kubernetes service within cluster):
# SQL_HOST=balancer-postgres-rw
# SQL_HOST=balancer-postgres-rw.balancer.svc.cluster.local
# CloudNativePG (primary — Philthy Civic Cloud shared cluster):
# SQL_HOST=shared-cluster-rw.cloudnative-pg.svc.cluster.local
# SQL_DATABASE=balancer
# (SSL typically not required within cluster)
#
# AWS RDS (External database):
# AWS RDS (legacy — for migration contexts):
# SQL_HOST=balancer-db.xxxxx.us-east-1.rds.amazonaws.com
# SQL_DATABASE=balancer_dev
# (SSL typically required - set SQL_SSL_MODE if needed)
#
# Local development:
Expand Down
9 changes: 7 additions & 2 deletions deploy/manifests/balancer/base/secret.template.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,9 @@
# repository. Secrets should be created in each target cluster using cluster-specific
# tools (e.g., SealedSecrets in the cfp-sandbox-cluster).
#
# CloudNativePG is the primary database target (Philthy Civic Cloud shared cluster).
# AWS RDS was used previously and may still be referenced for migration contexts.
#
apiVersion: v1
kind: Secret
metadata:
Expand All @@ -20,8 +23,10 @@ stringData:
REACT_APP_API_BASE_URL: https://balancer.sandbox.k8s.phl.io/
SECRET_KEY: randomly_generated_key_ere
SQL_ENGINE: django.db.backends.postgresql
SQL_HOST: sql_host_here
# CloudNativePG (primary): shared-cluster-rw.cloudnative-pg.svc.cluster.local
# AWS RDS (legacy): balancer-db.xxxxx.us-east-1.rds.amazonaws.com
SQL_HOST: shared-cluster-rw.cloudnative-pg.svc.cluster.local
SQL_PORT: '5432'
SQL_DATABASE: balancer_dev
SQL_DATABASE: balancer
SQL_USER: balancer
SQL_PASSWORD: sql_password_here
4 changes: 2 additions & 2 deletions frontend/src/components/Header/MdNavBar.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -89,10 +89,10 @@ const MdNavBar = (props: LoginFormProps) => {
</li>
<li className="border-b border-gray-300 p-4">
<Link
to="/login"
to="/"
className="mr-9 text-black hover:border-b-2 hover:border-blue-600 hover:text-black hover:no-underline"
>
Medical Suggester
Medication Suggester
</Link>
</li>
<li className="border-b border-gray-300 p-4">
Expand Down
1 change: 1 addition & 0 deletions frontend/src/pages/PatientManager/NewPatientForm.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -460,6 +460,7 @@ const NewPatientForm = ({
id="suicide"
name="suicide"
type="radio"
value="No"
checked={newPatientInfo.Suicide === "No"}
onChange={(e) => handleRadioChange(e, "Suicide")}
className="w-4 h-4 text-indigo-600 border-gray-300 focus:ring-indigo-600"
Expand Down
153 changes: 153 additions & 0 deletions server/api/views/assistant/agentic_loop.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
import json
import logging

from api.views.assistant.assistant_types import (
AgentResult,
ToolCallExecution,
ToolCallStatus,
TokenUsage,
)

logger = logging.getLogger(__name__)


def run_agentic_loop(
response, client, model_defaults: dict, tools: list, user
) -> AgentResult:

# Every tool call the agentic loop made before exiting
agentic_loop_tool_call_executions= []
# Token usage for every responses.create call, one entry per iteration
agentic_loop_token_usage: list[TokenUsage] = []

# TODO: Cap the number of iterations — a model that keeps calling tools never exits, hanging the request and adding OpenAI cost
while True:
# At the top of the body, so the initial response and the terminal iteration are each counted exactly once
# get_token_usage never raises: it runs on the web request path, so an unrecognized usage shape must not fail a user's request.
agentic_loop_token_usage.append(get_token_usage(response))

# TODO: Add a schema function to ToolCallExecution
# user is threaded through so tools that need it get it at dispatch time
tool_output_schemas, tool_call_executions = handle_tool_calls(response, tools, user)

# TODO: Rewrite to .append every iteration's list of tools or decide whether
# to add iteration integer to ToolCallExecution to split the flat tool_calls into iterations

# TODO: Add a data type to contain each iteration's parameters, response id,
# token usage, and tool calls or output text from the client response and a function
# for the output text and corresponding response id

# .extend splices every iteration's list of tools into one list
agentic_loop_tool_call_executions.extend(tool_call_executions)

# Exit agentic loop when model response doesn't contain any tool calls
if not tool_output_schemas:
return AgentResult(
output_text=response.output_text,
response_id=response.id,
tool_calls=agentic_loop_tool_call_executions,
token_usages=agentic_loop_token_usage,
)

# TODO: Add error handling to collect partial AgentResult tool calls and token usage —
# decide first whether AgentResult describes only successful runs or whatever happened
response = client.responses.create(
input=tool_output_schemas,
previous_response_id=response.id,
**model_defaults,
)

def get_token_usage(response) -> TokenUsage:
"""Token usage for one response"""

# Guard the whole chain when usage is missing because this also runs on the web request path
# Field names from openai ResponseUsage, checked against 2.29.0. requirements.txt doesn't pin openai,
# and nothing here checks types, so an SDK upgrade could blank or change these silently

usage = getattr(response, "usage", None)
input_details = getattr(usage, "input_tokens_details", None)
output_details = getattr(usage, "output_tokens_details", None)

return TokenUsage(
input_tokens=getattr(usage, "input_tokens", None),
cached_input_tokens=getattr(input_details, "cached_tokens", None),
output_tokens=getattr(usage, "output_tokens", None),
reasoning_output_tokens=getattr(output_details, "reasoning_tokens", None),
)


def handle_tool_calls(
response, tools: list, user
) -> tuple[list[dict], list[ToolCallExecution]]:

# Index the tools by name so a model-supplied call name can be looked up. .get()
# returns None for an unknown name, handled explicitly below.
tools_by_name = {tool.name: tool for tool in tools}

tool_output_schemas = []
tool_call_executions: list[ToolCallExecution] = []

for response_item in response.output:
if response_item.type == "reasoning":
#logger.info(f"Reasoning step: {response_item.summary}")
pass

elif response_item.type == "function_call":

tool_output, tool_call_execution = _execute_function_call(response_item, tools_by_name, user)

tool_output_schemas.append(
{
"type": "function_call_output",
"call_id": response_item.call_id,
"output": tool_output,
}
)

tool_call_executions.append(tool_call_execution)


return tool_output_schemas, tool_call_executions


def _execute_function_call(
response_item, tools_by_name: dict, user
) -> tuple[str, ToolCallExecution]:

target_tool = tools_by_name.get(response_item.name)

# Parsed below; stays None if the model's argument JSON can't be parsed,
# so a FAILED record still reports whatever we managed to read.
arguments = None

if target_tool is None:
msg = f"ERROR - No tool registered for function call: {response_item.name}"
logger.error(msg)
return msg, ToolCallExecution(
name=response_item.name,
status=ToolCallStatus.UNREGISTERED,
error=msg,
)

try:
arguments = json.loads(response_item.arguments)
logger.info(
f"Invoking tool: {response_item.name} with arguments: {arguments}"
)
tool_output = target_tool.run(user=user, **arguments)
logger.info(f"Tool {response_item.name} completed successfully")
return tool_output, ToolCallExecution(
name=response_item.name,
status=ToolCallStatus.OK,
arguments=arguments,
output=tool_output,
)
except Exception as e:
msg = f"Error executing function call: {response_item.name}: {e}"
logger.error(msg, exc_info=True)
return msg, ToolCallExecution(
name=response_item.name,
status=ToolCallStatus.FAILED,
arguments=arguments,
error=str(e),
)
12 changes: 12 additions & 0 deletions server/api/views/assistant/assistant_prompts.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,15 @@

# TODO: Replace the {name}/{page_number} citation template with a filled-in example,
# e.g. [Name advancespharmaco.pdf, Page 9], and require exactly one page per citation.
# Note both known importers pass their string through verbatim (no .format() reads the braces)
# The only thing interpreting the braces is the model
# This and the UUID in search_tool.py block the eval's citation scoring
# Citations are unparseable until both land, which blocks the eval's scoring layer: citation
# accuracy is the cheapest real signal available, and a parser written before these two
# fixes would measure prompt drift rather than accuracy.

# TODO: When ask_database is registered again, mention it here — the prompt names only
# search_documents and says to "ALWAYS use" it first, steering the model away from ask_database
INSTRUCTIONS = """
You are an AI assistant that helps users find and understand information about bipolar disorder
from your internal library of bipolar disorder research sources using semantic search.
Expand Down
71 changes: 26 additions & 45 deletions server/api/views/assistant/assistant_services.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,70 +3,51 @@

from openai import OpenAI

from .assistant_prompts import INSTRUCTIONS
from .tool_services import (
SEARCH_TOOLS_SCHEMA,
make_search_tool_mapping,
handle_tool_calls_with_reasoning,
)
from api.views.assistant.assistant_prompts import INSTRUCTIONS
from api.views.assistant.tool_services import TOOLS
from api.views.assistant.assistant_types import AgentResult
from api.views.assistant.agentic_loop import run_agentic_loop

logger = logging.getLogger(__name__)

# Module-level so eval_assistant.py can import it and log which model the run used
MODEL_NAME = "gpt-5-nano"


def run_assistant(
message: str,
user,
message: str,
previous_response_id: str | None = None,
) -> tuple[str, str]:
"""Wire together the OpenAI client, retrieval, and the agentic reasoning loop.

Parameters
----------
message : str
The user's input message.
user : User
The Django user object used for document access control in search_documents.
previous_response_id : str | None
ID of a prior response for multi-turn conversation continuity.

Returns
-------
tuple[str, str]
(final_response_output_text, final_response_id)
"""
# TODO: Track total duration, cost metrics, and tool_calls_made count
# and return them from run_assistant for use in eval_assistant.py CSV output
) -> AgentResult:

client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))

MODEL_DEFAULTS = {
"instructions": INSTRUCTIONS,
"model": "gpt-5-nano", # 400,000 token context window
# A summary of the reasoning performed by the model. This can be useful for debugging and understanding the model's reasoning process.
"model": MODEL_NAME,
# TODO: Flip "summary" to "auto" once this org is confirmed verified with OpenAI
"reasoning": {"effort": "low", "summary": None},
"tools": SEARCH_TOOLS_SCHEMA,
"tools": [tool.schema() for tool in TOOLS],
}

# TOOLS_SCHEMA tells the model what tools exist and what arguments to generate.
# tool_mapping wires those tool names to the Python functions that execute them.
# They are separate because the model generates arguments (schema concern) but
# cannot supply request-time values like user (mapping concern).
tool_mapping = make_search_tool_mapping(user)

if not previous_response_id:
response = client.responses.create(
input=[
{"type": "message", "role": "user", "content": str(message)}
],
**MODEL_DEFAULTS,
)
else:
response = client.responses.create(
if previous_response_id:
initial_response = client.responses.create(
input=[
{"type": "message", "role": "user", "content": str(message)}
],
previous_response_id=str(previous_response_id),
**MODEL_DEFAULTS,
)

return handle_tool_calls_with_reasoning(response, client, MODEL_DEFAULTS, tool_mapping)
# search_documents needs the request user for document access control
return run_agentic_loop(initial_response, client, MODEL_DEFAULTS, TOOLS, user)

initial_response = client.responses.create(
input=[
{"type": "message", "role": "user", "content": str(message)}
],
**MODEL_DEFAULTS,
)

# search_documents needs the request user for document access control
return run_agentic_loop(initial_response, client, MODEL_DEFAULTS, TOOLS, user)
Loading
Loading