Automation scripts for testing Deepgram services running on Amazon SageMaker as an "Endpoint" resource.
skills/deepgram-sagemaker/ is an installable agent
skill that walks a customer from AWS Marketplace subscription to a tested
SageMaker endpoint, running every AWS step through a deterministic script
instead of hand-typed console or CLI work. It follows the open
SKILL.md format, so it works in Claude Code, Codex,
Cursor and other agents that read skills.
Install into a project:
npx skills add deepgram-devs/dg-sagemaker # any SKILL.md-aware agent
# Claude Code plugin:
/plugin marketplace add deepgram-devs/dg-sagemaker
/plugin install deepgram-sagemaker@deepgramThen ask the agent, e.g. "set up Deepgram Nova-3 streaming on SageMaker in
us-east-2". It needs AWS credentials for the target account and
uv (each script declares its own dependencies).
The agent confirms with you before anything that costs money (subscribing,
creating an endpoint) or deletes resources, recommends an instance pool
rather than a single type, and refuses asynchronous endpoints, which are
temporarily not supported for Marketplace-hosted Deepgram (ask a Deepgram
representative if you need them). For capacity planning it will tell you to
measure concurrency on your own endpoint — or ask a Deepgram representative —
rather than quote a number.
The scripts also work on their own (uv run skills/deepgram-sagemaker/scripts/<script>.py --help):
preflight.py— credentials, region, tool versions, and a permission smoke test per phaselist_products.py— Deepgram's SageMaker listings, this account's subscription state (ACTIVE agreements only) and which offers it can accept (public / private / AWS Marketplace Field Demonstration Program)subscribe.py— the Marketplace Agreement API subscribe flow; quotes without--accept, subscribes with it; prefers a Field Demonstration Program offer when the account sees one (--no-fdpfor the public offer)resolve_model_package_arn.py— product + version → per-region ModelPackage ARN, recommended and supported instance typescheck_quota.py— per-type endpoint quota, current usage, in-flight requests;--request Nopens an increasecreate_execution_role.py— idempotent SageMaker execution role (+ S3 grant for async)deploy_endpoint.py— Model + EndpointConfig + Endpoint with the required settings baked in; waits and diagnoses failuresendpoint_status.py— status, variant, log tail, named cause + next stepinvoke_test.py— one real streaming or synchronous request with the right path and params; explains the classic 400sconfigure_autoscaling.py— target-tracking auto-scaling for real-time endpointsupdate_endpoint.py— in-place update (instance type/count, AMI, env, model version) with rollbackteardown_endpoint.py— delete endpoint + config + model, then verify nothing is left
Reference material the skill reads lives in skills/deepgram-sagemaker/references/
(products.json is the machine-readable catalog of listings, API paths, required parameters and instance types).
See js-stt/README.md for setup and usage. Built on the AWS SDK HTTP/2 bidirectional streaming client (@aws-sdk/client-sagemaker-runtime-http2); configuration (region, endpoint name, input file, query string) is edited inline at the top of each script.
Scripts:
stt.file.ts— streams a WAV file to a bidirectional streaming endpoint, chunking the file with keepalivesstt.microphone.ts— captures live microphone input and streams it to a bidirectional streaming endpointstress-stt.ts— fires N parallelstt.file.tsinvocations and reports success/failure counts and timing
See python-stt/README.md for full setup and usage.
Scripts:
stt_microphone_stress.py— streams live microphone audio; supports multiple simultaneous connectionsstt_wav_stress.pystream— streams a WAV file at real-time pace; repeatable load testing without a microphonestt_wav_stress.pybatch— posts WAV files via HTTP with configurable concurrency; reports latency and throughputstt_wav_async.py— transcribes a WAV file (up to 1 GiB) via the SageMakerInvokeEndpointAsyncAPI with S3 input/output; suits long-form audio beyond the synchronous invoke limit, with configurable concurrency and a latency/throughput summary
End-to-end correctness gates (python-stt/e2e/) — wrap the stress scripts and score each connection's transcript against a known reference (spacewalk.wav) via Word Error Rate; intended as the promotion gate before an endpoint goes live:
e2e/e2e_test_streaming.py— drivesstt_wav_stress.py streamthrough ~10 scenarios (basic short/long-form, sustained + ramped concurrency, the major feature flags, an adversarial WebSocket-close path) and checks each connection's combined final transcript by WERe2e/e2e_test_batch.py—--mode sync(25 s sample viainvoke_endpoint, ≤ 25 MB) or--mode async(~15 min / ~76 MB viainvoke_endpoint_async+ S3, incl. summarize); validates every returned transcript by WER
See java/README.md for an index of Java projects.
java/stt/aws-sdk— WAV streaming load test built directly on AWS SDK v2 HTTP/2 bidi streamingjava/stt/deepgram-sdk— same load test, via the Deepgram Java SDK + SageMaker transport
TBD
See python-tts/README.md for full setup and usage.
Scripts:
tts_stress.py— streams text phrases to multiple simultaneous bidirectional connections; plays audio from one selectable connection
End-to-end correctness gates (python-tts/e2e/) — validate the synthesized audio itself (non-empty, correct container/codec, non-silent, requested sample rate, speed→duration), so no second transcription endpoint is required:
e2e/e2e_test_batch.py— synchronousinvoke_endpointagainst/v1/speak; carries the full parameter matrix (model/encoding/sample_rate/bit_rate/container/speed, inline IPA override, 2000-char limit)e2e/e2e_test_streaming.py— websocket/v1/speak; the streaming-only behaviors (Speak→audio,Flush→Flushed, sustained concurrency, streaming encodings, voice/speed)
See python-flux-tts/README.md for full setup and usage.
Flux TTS uses /v2/speak and a turn-based protocol (Speak … Flush), so it has
its own client rather than sharing the Aura-2 one. Requires ml.g6.2xlarge or
newer — it does not run on g5 or g4dn.
Scripts:
flux_tts_client.py— shared client for both surfaces:FluxTtsStream(websocket/v2/speakover SageMaker bidirectional streaming) andinvoke_batch()(POST /invocations)
End-to-end correctness gates (python-flux-tts/e2e/) — both transports are served by the same endpoint, so one deployment covers both drivers:
e2e/e2e_test_batch.py—POST /invocations→/v2/speak; audio validity,container=wav, speed→duration, and negative controls (unknown param, Aura model rejected)e2e/e2e_test_streaming.py— websocket/v2/speak; the turn-based behaviors (per-turnSpeechMetadataaccounting, multi-turn, incrementalSpeak,Interrupt, mid-streamConfigure{speed}, concurrency)
See python-flux/README.md for full setup and usage.
Scripts:
flux_stress.pyfile— streams a WAV file to multiple Flux connections at real-time paceflux_stress.pymicrophone— streams live microphone audio to multiple Flux connectionsflux_stress.pylist-endpoints— lists available SageMaker endpoints in the target region
End-to-end correctness gate (python-flux/e2e/) — Flux is streaming-only (/v2/listen), so a single driver covers it:
e2e/e2e_test_streaming.py— drivesflux_stress.py filethrough basic / concurrency / connection-param / multilingual / in-band-control / negative scenarios, scoring each connection's combinedEndOfTurntranscript against the reference by WER