Skip to content

fix(app/bedrock-smartllm-backend): start prometheus_client server on 9100 so ServiceMonitor scrape works - #57

Merged
JLCode-tech merged 1 commit into
release/2.2from
fix/bedrock-shim-metrics-port
May 3, 2026
Merged

JLCode-tech merged 1 commit into
release/2.2from
fix/bedrock-shim-metrics-port

Conversation

@JLCode-tech

Copy link
Copy Markdown
Owner

Summary

The bedrock-shim's Service has a metrics named port at 9100, and the ServiceMonitor scrapes :metrics (port name → 9100). But nothing was listening on 9100 — /metrics was only mounted on the FastAPI app at CHAT_PORT (8080). Result: Prometheus reported up=0 for all 3 backend shims (nova-lite, nova-micro, claude-3-haiku) despite the pods being healthy and serving chat traffic correctly.

This forced a cluster-side patch on fresh deploys: redirect the ServiceMonitor's port: metrics → port: chat so it scrapes 8080 instead. Documented in agent memory project_aws_syd_test_tmm_kernelmode_break.md § "Update 2026-05-01".

Fix

Add prometheus_client.start_http_server(METRICS_PORT) (default 9100) before uvicorn.run() in the shim's __main__ block. This binds port 9100 with the same metric registry as the FastAPI /metrics endpoint. ServiceMonitor scrape now succeeds without any cluster-side patch.

The FastAPI /metrics route stays as a harmless duplicate so ad-hoc curl on either port keeps working.

Verification

After this PR is in the catalog and a fresh deploy runs:

  • kubectl -n monitoring exec deploy/prometheus -- wget -qO- 'http://localhost:9090/api/v1/query?query=up{job=~"nova-lite|nova-micro|claude-3-haiku"}' should show value: "1" for all 3 (currently "0" without the cluster-side patch)
  • vllm:num_requests_running, vllm:gpu_cache_usage_perc, etc series should populate without delay

Related

  • bnk-forge#85 — uses these metrics in the new Backend Health panel via direct K8s API server proxy scrape; works around the same port-mismatch from the forge backend side. With this PR merged, the ServiceMonitor path also works → real Prometheus dashboards become useful too.

…9100 so ServiceMonitor scrape works

The shim previously mounted /metrics only on the FastAPI app at CHAT_PORT
(8080). The accompanying Service exposes port 9100 named "metrics" and
the ServiceMonitor scrapes :metrics, but nothing was listening on 9100 →
Prometheus reported up=0 for all 3 backends despite the pods being
healthy.

Adding prometheus_client.start_http_server(METRICS_PORT) before
uvicorn.run binds 9100 with the same metric registry. ServiceMonitor
scrape now succeeds without needing any cluster-side port redirect.

Keeps the FastAPI /metrics route too — harmless duplicate that lets
ad-hoc curl on either port keep working.
@JLCode-tech
JLCode-tech merged commit 94c5e64 into release/2.2 May 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant