fix(app/bedrock-smartllm-backend): start prometheus_client server on 9100 so ServiceMonitor scrape works - #57
Merged
Conversation
…9100 so ServiceMonitor scrape works The shim previously mounted /metrics only on the FastAPI app at CHAT_PORT (8080). The accompanying Service exposes port 9100 named "metrics" and the ServiceMonitor scrapes :metrics, but nothing was listening on 9100 → Prometheus reported up=0 for all 3 backends despite the pods being healthy. Adding prometheus_client.start_http_server(METRICS_PORT) before uvicorn.run binds 9100 with the same metric registry. ServiceMonitor scrape now succeeds without needing any cluster-side port redirect. Keeps the FastAPI /metrics route too — harmless duplicate that lets ad-hoc curl on either port keep working.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The bedrock-shim's Service has a
metricsnamed port at 9100, and the ServiceMonitor scrapes:metrics(port name → 9100). But nothing was listening on 9100 —/metricswas only mounted on the FastAPI app atCHAT_PORT(8080). Result: Prometheus reportedup=0for all 3 backend shims (nova-lite,nova-micro,claude-3-haiku) despite the pods being healthy and serving chat traffic correctly.This forced a cluster-side patch on fresh deploys: redirect the ServiceMonitor's
port: metrics→port: chatso it scrapes 8080 instead. Documented in agent memoryproject_aws_syd_test_tmm_kernelmode_break.md§ "Update 2026-05-01".Fix
Add
prometheus_client.start_http_server(METRICS_PORT)(default 9100) beforeuvicorn.run()in the shim's__main__block. This binds port 9100 with the same metric registry as the FastAPI/metricsendpoint. ServiceMonitor scrape now succeeds without any cluster-side patch.The FastAPI
/metricsroute stays as a harmless duplicate so ad-hoccurlon either port keeps working.Verification
After this PR is in the catalog and a fresh deploy runs:
kubectl -n monitoring exec deploy/prometheus -- wget -qO- 'http://localhost:9090/api/v1/query?query=up{job=~"nova-lite|nova-micro|claude-3-haiku"}'should showvalue: "1"for all 3 (currently"0"without the cluster-side patch)vllm:num_requests_running,vllm:gpu_cache_usage_perc, etc series should populate without delayRelated