A vendor-neutral reference solution for automating clinical medical coding — mapping
free-text "verbatims" from clinical trial data (e.g. "migrane", "Hypothyroid") to
controlled-vocabulary codes (MedDRA for adverse events, WHODrug for medications) —
using a Step Functions state machine with a declarative Amazon Bedrock AgentCore
Harness step for the cases that need semantic reasoning.
The study/site/subject identifiers and free-text verbatims in data/ are fabricated; the
dictionary terms and codes are a de minimis illustrative set based on the shape of
real MedDRA (MSSO/IFPMA) and WHODrug (Uppsala Monitoring Centre) records — not a
redistributable dictionary extract. No customer or patient data is included.
- Node.js 24+ and npm.
- The AWS CLI v2, configured with credentials for the target account (
aws configureor an SSO profile). - The AWS CDK CLI — installed via
npm installbelow (aws-cdkis a dev dependency), or globally withnpm install -g aws-cdk. - Your target account/region bootstrapped for CDK (
npx cdk bootstrap), if this is the first CDK deploy there. jqanduuidgen— used only by the shell snippets in the testing section below, not by the deploy or seed steps.uuidgenships with macOS and most Linux distros; on a minimal Amazon Linux 2023 host neither is present, so install both withsudo dnf install -y jq util-linux(jqviabrew install jq/apt install jqelsewhere).- Model access enabled in the Amazon Bedrock console for Titan Text Embeddings V2 (
amazon.titan-embed-text-v2:0, used by the seed script and the dictionary-search tool) and Claude Sonnet (the harness'sus.anthropic.claude-sonnet-4-6inference profile — seelib/constructs/coding-harness.ts), in the region you deploy to.
Cheapest-first escalation, with the workflow — not the model — owning every routing decision:
- Deterministic direct lookup (Lambda, no LLM). A block-list guard runs first:
an exact hit on
terms_not_to_autocodemeans the term is too ambiguous to ever auto-assign, and the row stays open for a human. Then an exact/synonym match againstdictionary_terms+synonym_list— accepted only if it resolves to exactly one distinct code. Most verbatims are coded here, cheaply and reproducibly, with a score of 1.00. - Agentic semantic search and adjudication (AgentCore Harness), only for
verbatims the deterministic path missed. The harness is declared directly in
the Step Functions task state — model, system prompt, and two tools exposed
through an AgentCore Gateway, with no container to build or operate:
search_dictionary(pgvector cosine-similarity search overdictionary_terms) andget_study_info(the study's free-text metadata description). The second tool exists because similarity search can return several clinically plausible candidates it cannot separate on text alone — a human coder breaks that tie using study context, and this gives the agent the same context. Choosing among close candidates is the part that resists being written as code, which is why this step is an agent rather than another Lambda. - Score-threshold routing (Choice state, configuration not model judgment): high confidence autocodes, medium confidence routes to a human review queue, low confidence leaves the term open.
- Write-back (Lambda). Persists the outcome onto the same
study_termsrow in place — status, the winning code, its full dictionary hierarchy, and astatus_changed_byaudit marker distinguishing machine coding from human action.
Step Functions owns sequencing, retries, and the audit trail (the execution history
is the audit trail — no separate logging needed). The agent owns reasoning within
one bounded step and cannot skip a gate or reorder the workflow. See
state-machine/coding-workflow.asl.yaml for the full annotated definition.
One Aurora PostgreSQL (Serverless v2) cluster, five tables, accessed entirely through the RDS Data API (no VPC attachment needed by the Lambdas):
| Table | Role |
|---|---|
dictionary_terms |
The target vocabulary (MedDRA LLT/PT or WHODrug trade/generic names), each row carrying its full hierarchy (JSONB) and a Titan v2 embedding (VECTOR(1024), HNSW index) of its normalized text. |
synonym_list |
Curated verbatim → code shortcuts that widen "exact match" beyond literal dictionary text (e.g. "Hypothyroid" → Hypothyroidism / 10021114). |
terms_not_to_autocode |
The safety block-list, checked first — a hit is a deliberate "never auto-assign," not "no match found." |
study_terms |
The work queue: one row per verbatim, holding both the input (verbatim, dictionary, version, study) and the coding output (status, code, hierarchy, score) on the same row. |
study_metadata |
What each trial is about, in prose — three columns (id, study name, description). Read by get_study_info so the agent can break ties between candidate terms. Joined to study_terms.source_study by name. |
See lib/constructs/coding-database.ts for the CDK-managed cluster and
scripts/seed.ts for the schema DDL and sample data loader.
lib/
├── coding-stack.ts # top-level CDK stack
└── constructs/
├── coding-database.ts # Aurora Serverless v2 + pgvector
├── coding-lambdas.ts # checkDirect, dictionarySearch, writeBack
├── coding-gateway.ts # AgentCore Gateway fronting the search tool
├── coding-harness.ts # AgentCore Harness resource
└── coding-state-machine.ts # Step Functions state machine + role
lambda/
├── checkDirect/ # State 1 — block-list + exact/synonym match
├── tools/dictionarySearch/ # Gateway tool — pgvector semantic search
├── tools/studyInfo/ # Gateway tool — study metadata lookup
├── writeBack/ # State 4 — persist outcome to study_terms
└── shared/dataApi.ts # RDS Data API helper
state-machine/
└── coding-workflow.asl.yaml # the state machine (JSONata)
scripts/
└── seed.ts # creates schema, loads data/*.csv, embeds
test/
├── checkDirect.test.ts # deterministic-path decision logic
└── stateMachine.test.ts # ASL wiring + agent-reply parsing contract
data/
├── dictionary_terms_meddra.csv # (1) MedDRA target vocabulary
├── dictionary_terms_whodrug.csv # (1) WHODrug target vocabulary
├── synonym_list.csv # (2) curated verbatim -> code
├── terms_not_to_autocode.csv # (3) block-list
├── study_terms_input.csv # (4) input verbatims to code
├── study_metadata.csv # (5) study context for disambiguation
├── golden_walkthrough.csv # (6) 12 traced cases A-L (the tests below)
└── generate_sample_data.py # regenerates every CSV above
npm install
npm run build
npm test # unit tests: no AWS calls, no deployed stack needed
npx cdk deployNote the stack outputs — you'll need them below:
MedicalCodingAgentCoreDemo.DbClusterArn = arn:aws:rds:...:cluster:...
MedicalCodingAgentCoreDemo.DbName = coding
MedicalCodingAgentCoreDemo.DbSecretArn = arn:aws:secretsmanager:...
MedicalCodingAgentCoreDemo.StateMachineArn = arn:aws:states:...:stateMachine:MedicalCodingWorkflow
Creates the five tables and loads data/*.csv, generating a Titan v2 embedding for
every dictionary row along the way:
npm run seed
# equivalent to: npx tsx scripts/seed.ts --stack MedicalCodingAgentCoreDemoRe-run any time — it drops and recreates the tables, so it's always safe to reseed before a test pass. The script prints a row-count verification at the end; if Aurora had auto-paused, the first call may take ~15-30s to resume (the script retries automatically).
data/golden_walkthrough.csv traces one input verbatim through every branch the
state machine can take. After seeding, study_terms has twelve open rows — one per
case below. Each case starts a Step Functions execution and checks both the
execution output and the persisted database row.
| Case | Study | Verbatim | Path exercised | Expected outcome |
|---|---|---|---|---|
| A | ONCO-2024-01 | Dislocated shoulder |
exact dictionary match | autocoded, score 1.00 |
| B | ONCO-2024-01 | Hypothyroid |
synonym match | autocoded, score 1.00 |
| C | ONCO-2024-01 | HAEMORRHAGE |
block-list hit | open (safety block) |
| D | ONCO-2024-01 | migrane |
agent, high confidence | autocoded |
| E | ONCO-2024-01 | flutters in my chest sometimes |
agent, medium confidence | approval_required |
| F | ONCO-2024-01 | asdfghjkl |
agent, no confident match | open |
| G | ONCO-2024-01 | ibuprofen |
exact match (WHODrug) | autocoded, score 1.00 |
| H | ONCO-2024-01 | FULVESTRANT |
exact match, derivation_only |
autocoded, only derivation/hierarchy written, no new dict_term |
| I | GI-2025-03 | nexiuum |
agent, near-miss trade names split by study context | autocoded → NEXIUM (esomeprazole), not NEXIM (tranexamic acid) |
| J | ONCO-2024-01 | cramps |
agent, study context decides | approval_required → Muscle spasms |
| K | GI-2025-03 | cramps |
agent, same verbatim, other study | approval_required → Abdominal pain |
| L | CARD-2025-02 | heart races when I stand up |
agent, competing rhythm terms | approval_required → Tachycardia |
J and K are the same verbatim — cramps — in two different studies. Similarity
search returns the same candidate set for both (Muscle spasms, Myalgia,
Abdominal pain), and nothing in the verbatim itself can separate them. The only
differentiator is what get_study_info returns:
- ONCO-2024-01 is an endocrine-therapy breast cancer study that actively
monitors muscle cramps and spasms, and explicitly does not treat GI events as an
endpoint →
Muscle spasms(10028334). - GI-2025-03 is a reflux study where abdominal cramping is an expected adverse
event of special interest and musculoskeletal events are not an endpoint →
Abdominal pain(10000081).
Identical input, different code, decided entirely by study context — the same way a
human coder who knows the protocol would decide it. Case I is the drug-side version
of the same idea: NEXIM (tranexamic acid, an antifibrinolytic) and NEXIUM
(esomeprazole, a PPI) are both lexically near-identical to the misspelled nexiuum,
and only the study's expected concomitant medications tell them apart.
Assert on the code, not the status, for J and K. The selected term is stable
across runs — five consecutive executions of case J returned Muscle spasms every
time. The confidence is not: the same input scored 0.85, 0.85, 0.88, 0.88 and 0.92
across those runs, which straddles the 0.90 autocode gate, so J legitimately lands in
either autocoded or approval_required depending on the run. That is inherent to a
model-authored score, not a defect — but it means a verbatim whose confidence sits
near a threshold is not a deterministic fixture. If you need a reproducible assertion
for CI, compare dict_term_code; treat the status as a band.
The score the agent returns — and therefore the value ScoreThreshold routes on —
is the agent's own coding confidence, not the retrieval cosine similarity that
search_dictionary reports. The two are different measurements and are not
interchangeable:
| Verbatim | Retrieval cosine | Coding confidence | Why they differ |
|---|---|---|---|
migrane |
0.37 | 0.99 | A misspelling is lexically distant from its correct term even when the coding decision is unambiguous. |
flutters in my chest sometimes |
0.65 | 0.88 | Lay phrasing overlaps several rhythm terms; the decision is genuinely less certain than D. |
Routing on raw cosine similarity inverts these two cases, which is why the prompt asks for a confidence judgment instead. Retrieval correctness (is the right candidate in the list?) and routing confidence (how sure are we of the choice?) are separate concerns, and only the first is a property of the embedding.
This means the routing signal is model-authored. That is deliberate — it is the only signal available that tracks coding difficulty rather than string distance — but it is defended by controls that are not model-authored:
- the block-list runs before any model call and cannot be overridden by the agent;
WriteBackverifies the chosen code actually exists in the target dictionary and version (verifyCodeExists), so a fabricated or study-description-derived code is rejected rather than persisted;- the 0.90/0.70 gates live in the state machine definition, not in the prompt, so they are configuration a reviewer can audit and tune;
- anything below the autocode gate reaches a human.
The 0.90/0.70 values themselves remain illustrative starting points. Calibrate them against your own dictionary and verbatim distribution before relying on them.
export STACK=MedicalCodingAgentCoreDemo
export REGION=us-east-1 # or your deploy region
SM_ARN=$(aws cloudformation describe-stacks --stack-name "$STACK" --region "$REGION" \
--query "Stacks[0].Outputs[?OutputKey=='StateMachineArn'].OutputValue" --output text)
CLUSTER_ARN=$(aws cloudformation describe-stacks --stack-name "$STACK" --region "$REGION" \
--query "Stacks[0].Outputs[?OutputKey=='DbClusterArn'].OutputValue" --output text)
SECRET_ARN=$(aws cloudformation describe-stacks --stack-name "$STACK" --region "$REGION" \
--query "Stacks[0].Outputs[?OutputKey=='DbSecretArn'].OutputValue" --output text)
DB_NAME=$(aws cloudformation describe-stacks --stack-name "$STACK" --region "$REGION" \
--query "Stacks[0].Outputs[?OutputKey=='DbName'].OutputValue" --output text)
# 1. Look up the record_id for the case you want to run (e.g. case D, "migrane")
aws rds-data execute-statement --region "$REGION" \
--resource-arn "$CLUSTER_ARN" --secret-arn "$SECRET_ARN" --database "$DB_NAME" \
--sql "SELECT record_id, verbatim, encoding_dictionary, encoding_dictionary_version,
source_study, derivation_only
FROM study_terms WHERE verbatim = 'migrane'"
# 2. Start an execution with that row's fields as input
aws stepfunctions start-execution --region "$REGION" \
--state-machine-arn "$SM_ARN" \
--name "case-D-$(date +%s)" \
--input '{
"record_id": "<record_id from step 1>",
"verbatim": "migrane",
"encoding_dictionary": "MedDRA",
"encoding_dictionary_version": "v27.0",
"source_study": "ONCO-2024-01",
"derivation_only": false
}'
# -> returns an executionArn; save it
# 3. Check the response: execution status + output
aws stepfunctions describe-execution --region "$REGION" \
--execution-arn "<executionArn from step 2>" \
--query "{status: status, output: output}"
# 4. Check the response: the persisted row (the source of truth)
aws rds-data execute-statement --region "$REGION" \
--resource-arn "$CLUSTER_ARN" --secret-arn "$SECRET_ARN" --database "$DB_NAME" \
--sql "SELECT status, status_changed_by, dict_term, dict_term_code, derivation, match_score
FROM study_terms WHERE verbatim = 'migrane'"A SUCCEEDED execution with a populated status/dict_term_code in the DB confirms
the case coded as expected. For the block-list and no-match cases (C, F), the
execution still SUCCEEDED — an un-coded term is a valid business outcome the
workflow reaches deliberately (via MarkOpen), not a failure.
Loop over every open row seeded from study_terms_input.csv and start one
execution per row — this is the fastest way to exercise every branch in one pass:
RECORDS=$(aws rds-data execute-statement --region "$REGION" \
--resource-arn "$CLUSTER_ARN" --secret-arn "$SECRET_ARN" --database "$DB_NAME" \
--sql "SELECT record_id, verbatim, encoding_dictionary, encoding_dictionary_version,
source_study, derivation_only
FROM study_terms WHERE status = 'open' ORDER BY record_id" \
--query "records" --output json)
echo "$RECORDS" | jq -c '.[] | {
record_id: .[0].stringValue,
verbatim: .[1].stringValue,
encoding_dictionary: .[2].stringValue,
encoding_dictionary_version: .[3].stringValue,
source_study: .[4].stringValue,
derivation_only: .[5].booleanValue
}' | while read -r input; do
aws stepfunctions start-execution --region "$REGION" \
--state-machine-arn "$SM_ARN" \
--name "golden-$(uuidgen)" \
--input "$input"
done
# Poll until all executions leave RUNNING, then compare against the table above:
aws stepfunctions list-executions --region "$REGION" \
--state-machine-arn "$SM_ARN" --status-filter RUNNING --query "executions[].name"
# Full result set to diff against the expectations table:
aws rds-data execute-statement --region "$REGION" \
--resource-arn "$CLUSTER_ARN" --secret-arn "$SECRET_ARN" --database "$DB_NAME" \
--sql "SELECT verbatim, status, dict_term_code, match_score, status_changed_by
FROM study_terms ORDER BY record_id"If an execution's outcome doesn't match the table, read its history for the state that made the call:
aws stepfunctions get-execution-history --region "$REGION" \
--execution-arn "<executionArn>" \
--query "events[?type=='TaskSucceeded' || contains(type, 'Failed')].{type: type, id: id}"Two states worth checking directly when a semantic-search case (D, E, F, I) doesn't land where expected:
- CodingAgent's raw tool output — invoke either Gateway tool directly to see
what the agent sees, without going through the LLM:
# candidates + cosine scores aws lambda invoke --region "$REGION" \ --function-name coding-demo-tool-dictionary-search \ --payload '{"verbatim_term":"migrane","dictionary":"MedDRA","dictionary_version":"v27.0","top_k":5}' \ --cli-binary-format raw-in-base64-out /dev/stdout # study context used to break ties aws lambda invoke --region "$REGION" \ --function-name coding-demo-tool-study-info \ --payload '{"study_name":"ONCO-2024-01"}' \ --cli-binary-format raw-in-base64-out /dev/stdout
- The agent's parsed answer — in the execution history, the
invokeHarnessTaskSucceededevent's output contains the assistant's JSON text (term, code, hierarchy, and self-reported confidence score) thatScoreThresholdroutes on.
If a semantic-search case routes unexpectedly, check the confidence the agent returned, not the cosine score — see "How
scoreis defined" above. A case that lands inopenwith a correct top candidate usually means the agent reported low confidence, not that retrieval failed.
See CONTRIBUTING for more information.
This library is licensed under the MIT-0 License. See the LICENSE file.