AugmentWorks is regression testing for AI agents: it checks conversational responses, reported tool calls, and configured application state. It is for engineers and coding assistants who need a deterministic first assessment without treating a chatbot saying “done” as proof that state changed.
The CLI maps fixed assessment operations to an application on localhost, a private network, or another customer-configured endpoint using versioned YAML. The generic HTTP connector does not require a Python adapter or AugmentWorks target SDK, and no coding assistant is used in either runtime path.
| Path | What it is | Needs | Status |
|---|---|---|---|
Packaged demo |
Loopback-only synthetic refund target, isolated fixtures, and the real local runner/scorer. Shows a policy bug, then the same packet passing after the policy is enforced. | Node.js 20+ | In this @augmentworks/cli@0.3.7 package |
Local test --local |
Customer-executed scoring of a data-only packet against your configured target. Requires no AugmentWorks account and contacts no AugmentWorks service. | Node.js 20+, a connector YAML, and a target you are authorized to call. The bundled starter is an isolated synthetic fixture. | In this @augmentworks/cli@0.3.7 package |
Hosted test |
Outbound HTTPS relay assessment with a live dashboard. Browser approval does not start a run. Pending hosted judging is never a pass. | Invited workspace, login, and an authorized isolated synthetic or staging target while hosted real-data is release-disabled | In this @augmentworks/cli@0.3.7 package, including quoted --estimate / --max-credits |
Product site: https://augmentworks.ai.
Report schemas: schema --kind local-packet and schema --kind local-result.
Docs: quickstart,
local,
agent setup.
Do not invent popularity, certifications, endorsements, or AI search ranking. Local reports are unsigned, customer-executed evidence, not a certification, audit, or hosted evidence record.
| Identity | Current value |
|---|---|
This package (package.json) |
0.3.7 |
Executable npx pin (matches this tarball) |
@augmentworks/cli@0.3.7 |
| Hosted assessment protocol | aw-relay/0.3 |
| Immutable prior npm artifacts | @augmentworks/cli@0.3.6 (gitHead a9b927a2413305003f817c20e9c5df277512f83e, integrity sha512-idDM/kYfDCzDu+iaSqzZjEchyui9I8r1krUFdZ8BmVtZIye2D5WbHFpbYaNzuwjPvV7gCk1XixLwu1kuROXfwA==, published 2026-09-10T16:26:29.450Z; includes suite-selection capabilities), @augmentworks/cli@0.3.5 (gitHead 11570f6cf883ec6e6743e010c35134bb485234dd, integrity sha512-WyS9d2lSPhX26ONyxISbN9ncsDoR3JBQjkr6DLaJPl6IrksBg8zxpxawABb4iLZ4ugqlu7TA9J623nqua9dlwQ==, published 2026-09-09T02:56:49.964Z; omits suite-selection capabilities), @augmentworks/cli@0.3.4 (gitHead c3da8d92bdd3daa21e9e230ffc5d110b43adaa5f, integrity sha512-TLeAzDglZoGL6fWLxA9rIUwJd69NFqgmlONzU4uRmhDz4S31+dfSZpaj44ahb6lUjnmuoY6sDmfPStxFzInpVQ==, published 2026-09-08T06:38:24.847Z), and @augmentworks/cli@0.3.3 (gitHead 4a08ea0d352f2515e725cb9ca946807112422436). Do not overwrite or relabel any of them. |
| Hosted packet | support-refunds@0.1.0 |
| Local starter packet | support-refunds-starter@0.1.0 |
Executable npx examples pin 0.3.7, the identity of this package. This published-line package includes packaged demo, hosted --assessment / --profile, usage, billing, preview-mapping, probe, catalog list / show, selection compile with connector capabilities, suite validate / preview, test --suite, investigation inspect / fetch / export-regression, test --investigation, --estimate / --max-credits, run status / run wait / run report, compare / gate / baseline, AUGMENTWORKS_API_KEY mode, empty-directory own-target starters, native customer-suite upload translation, saved-suite v2, v2 gate receipts, durable execution IDs, cross-shard quote uniqueness, and non-destructive generated .env guidance. examples/ is still omitted from the npm tarball; starters ship under assets/. published_package_verified is published-line identity, not a live registry probe of this exact tarball. Independent inspection of @augmentworks/cli@0.3.7 (gitHead 876454878a31a897aec1d722fa838c9c0ecea2aa, published 2026-09-27T21:06:12.276Z, tarball SHA-256 37b448b4db2c195b5c361daab10604d57476fa4d9e441916f7bfe2054f81b313) is recorded in docs/feature-readiness/published-registry-evidence.json. The immutable 0.3.5 tarball omits suite-selection capabilities (AUG-82); independently inspected 0.3.6 includes that fix and does not include saved-suite /2. Website adoption of this exact 0.3.7 artifact is AUG-250. Do not overwrite or relabel npm 0.3.7, 0.3.6, 0.3.5, 0.3.4, or 0.3.3. npx --yes only skips the npm prompt; it is not a hosted spending ceiling. Do not run npx @augmentworks/cli@latest.
Prerequisites: Node.js 20 or newer. No API key, login, second terminal, Docker, database, or model.
From a clone of this repository after npm ci && npm run build:
node dist/index.js demoPublished npm:
npx --yes @augmentworks/cli@0.3.7 demoMachine-readable summary (not an AW-LOCAL-RESULT-1 report):
node dist/index.js demo --jsonThe command binds an isolated target to 127.0.0.1 on an OS-assigned port,
generates a per-run authentication value, and ignores caller CHATBOT_* /
AUGMENTWORKS_* environment variables, .env, and the working project's
YAML. It runs the real local scorer twice on the same data-only packet:
- Faulty implementation — refunds $80 even though
policy.maximum_refundis $50. Observationorder.statusbecomesrefunded. Expected assertionpolicy-limit-order-remains-paidfails. Underlying exit code10. - Corrected implementation — enforces the maximum. The order stays
paidandorder.refunded_amountstays0. Underlying exit code0.
A conversational “The synthetic order refund completed.” response is not
proof of correct state; the observer values are. The demo exits 0 only when
that fail-then-pass story and cleanup both succeed. Reports are written under
a fresh .augmentworks/demo/<id>/failing and passing directory. --open
opens HTML only if you pass it; default is no browser. After a hard kill,
cleanup may not run; the in-memory demo target vanishes with the process, but
a real application still needs a server-side fixture TTL.
Website adoption of independently inspected @augmentworks/cli@0.3.7 is
AUG-250. Do not use @latest, @0.3.6, or @0.3.5 for the saved-suite
catalog/selection path. The published tarball's embedded LAST_VERIFIED_*
still names 0.3.6; the inspection receipt is the later evidence file.
Prerequisites: Node.js 20 or newer, an invited AugmentWorks workspace, and an authorized, isolated synthetic or staging target while hosted real-data is release-disabled. Hosted access is not a public self-serve signup; do not assume a trial entitlement.
This @augmentworks/cli@0.3.7 package can log in, run --assessment, and
init writes augmentworks.assessment.yaml. From a clone after
npm ci && npm run build:
node dist/index.js init
node dist/index.js init --starter workflowinit writes augmentworks.yaml (or the path given to -c / --config),
augmentworks.assessment.yaml, packaged fixture servers, mapping fixtures,
and the referenced starter files for the selected own-target pattern
(response-quality / response-only JSON chat by default, or workflow /
stateful for support-refunds hooks). Response-only setup does not add unused
state hooks. It never overwrites an edited assessment or reference file unless
you pass --force. --force still never replaces an existing .env.
npm --yes only skips the npm installer prompt; it is not CLI spending consent.
npx --yes @augmentworks/cli@0.3.7 login
npx --yes @augmentworks/cli@0.3.7 init --agent
# This 0.3.7 package writes augmentworks.assessment.yaml and starter files.
# Edit .env for the authorized target. The packaged starter is an isolated synthetic fixture.
npx --yes @augmentworks/cli@0.3.7 doctor \
-c augmentworks.yaml
npx --yes @augmentworks/cli@0.3.7 test \
-c augmentworks.yaml \
--assessment ./augmentworks.assessment.yaml \
--profile quick \
--openlogin authorizes this terminal. Browser approval does not start an
assessment. This package's init creates the YAML, assessment file, starter
references, .env.example, a local .env, and optional repository guidance.
On POSIX systems, the CLI creates .env with mode 0600. doctor validates
the config, assessment, profile overrides, reference sizes, and wire bounds
without calling AugmentWorks or the target. Hosted test --assessment starts
one assessment, keeps this terminal online only for that run, and --open
opens its live dashboard.
There is no separate connect command. Keep the terminal open until the
assessment finishes. If grading is pending after target work, wait on the
original run; do not delete journals or blindly rerun the test command.
For an SSH or otherwise headless environment, use device authorization:
npx --yes @augmentworks/cli@0.3.7 login --deviceInteractive login can replace the one stored credential for this API origin
after it displays the selected workspace name and UUID. To pin a company
workspace (independently inspected @augmentworks/cli@0.3.7):
node dist/index.js login --workspace "$AUGMENTWORKS_WORKSPACE_ID"
node dist/index.js whoami --workspace "$AUGMENTWORKS_WORKSPACE_ID"--workspace and AUGMENTWORKS_WORKSPACE_ID must agree when both are set.
A machine key for another company fails WORKSPACE_MISMATCH before quote,
upload, or target work. test --local --workspace is rejected; the environment
variable is ignored in local mode and does not open a hosted session.
--profile still selects quick/full/smoke/release assessment profiles only.
usage reads the authenticated workspace ledger. It does not need target YAML,
a target API key, or a target server. It does not grant credits, reserve units,
create a run, or open checkout. Billing management stays in a signed-in browser
session with billing permission; this command only displays a server snapshot.
After npm ci && npm run build:
node dist/index.js usage
node dist/index.js usage --jsonHuman output identifies the workspace, shows available credits prominently,
then reserved and consumed credits with their ledger meanings. When the server
advertises subscriptions_v1, it also separates recurring, purchased, and
trial/promotional lots from the server grant balances and shows the current
service period, cancellation-at-period-end, monthly grant expiry, and next
payment action. Values are a server snapshot at asOf, not a guaranteed
future balance. Monthly grants expire at the server period end with no
initial rollover; purchased pack credits stay distinct. The CLI never
recomputes available credits by subtracting fields, never infers access from
a Stripe id or the local clock, and never treats a nullable expiry as
tomorrow. Login refresh, logout, and browser reauthentication do not imply a
new trial. Local demo, test --local, offline doctor, and schema
remain account-free and make no billing calls.
This package includes usage. Do not use @latest or the immutable npm 0.3.3
artifact for this command.
preview-mapping applies the production response extraction, allowlist,
redaction, and evidence serialization to a local synthetic JSON fixture. It
does not call the target, AugmentWorks, or a model, and it consumes no
credits. doctor still validates configuration only; use this inspector
before an assessment to see extracted versus missing fields and the exact
canonical evidence payload that would leave the machine.
After npm ci && npm run build:
node dist/index.js preview-mapping -c augmentworks.yaml --operation send --fixture ./fixtures/send-response.json
node dist/index.js preview-mapping -c augmentworks.yaml --operation send --fixture ./fixtures/send-response.json --jsonThe preview is for the supplied fixture only. It does not guarantee that future responses are secret-free. Do not pass production transcripts or live target output.
This package includes preview-mapping. Do not use @latest or the immutable
npm 0.3.3 artifact for this command.
Offline doctor and mapping preview do not prove network auth, selector
behavior, session continuity, or cleanup. probe is an explicit bounded
synthetic connection check. Without --yes it only prints the planned
operations, call count, time/byte limits, and possible synthetic side
effects. --yes then executes that plan against the configured target.
Doctor and init never start a probe. The probe does not contact AugmentWorks
or consume credits. Failures are integration diagnostics (connection,
timeout, authentication, selector, session, cleanup), not chatbot quality
verdicts.
After npm ci && npm run build:
node dist/index.js probe -c augmentworks.yaml
node dist/index.js probe -c augmentworks.yaml --json
node dist/index.js probe -c augmentworks.yaml --yesThis package includes probe. Do not use @latest or the immutable npm 0.3.3
artifact for this command.
billing retrieves the authenticated first-party billing page advertised by
billing_portal_link_v1 and prints or opens that URL. It does not create a
Stripe Customer, Checkout Session, purchase, refund, or subscription, and it
does not cancel, reactivate, or change payment methods. Recurring-plan
management stays on that first-party page after browser sign-in. Capability
subscriptions_v1 only controls whether usage/billing may show the server
subscription projection; it does not enable live $149 sales. The workspace id
in the URL is a navigation hint. Payment changes require a signed-in browser
session with owner/billing permission.
After npm ci && npm run build:
node dist/index.js billing
node dist/index.js billing --json
node dist/index.js billing --print--json and --print never open a browser. A failed GUI opener still prints
the safe URL and does not change credentials. If a hosted test is rejected for
insufficient credits, the CLI reports required vs available units when the
server supplied them, keeps the uncreated intent, and points at this page. Do
not wait in the terminal for a purchase. After fulfillment, run usage, then
start a new test explicitly with --max-credits. pendingCommerce on a usage
snapshot is processing metadata, not spendable credit. Pack prices belong to
the website catalog; this CLI does not market the $49 pack or the $149
monthly plan. If subscriptions_v1 is absent, recurring CTAs are omitted.
test --estimate compiles the same assessment that admission uses and calls
POST /v1/billing/quote. It does not create a run, reserve credits, call a
model, or contact the target. The server assessmentPlanHash is not the local
freeze hash. A quote is not a hold: another run may use credits before this
one starts.
Noninteractive hosted assessment tests require --max-credits N. --yes
skips the prompt but is not an unlimited budget. npx --yes is the npm
installer flag and is not CLI spending consent. Packet-only --packet
hosted tests keep aw-relay/0.1 and do not quote.
After npm ci && npm run build:
node dist/index.js test --assessment ./augmentworks.assessment.yaml --estimate
node dist/index.js test --assessment ./augmentworks.assessment.yaml --estimate --json
node dist/index.js test --assessment ./augmentworks.assessment.yaml --profile quick --max-credits 30 --yes
node dist/index.js run status <run-id>
node dist/index.js run wait <run-id>
node dist/index.js run report <run-id> --json
node dist/index.js compare --run <run-id> --baseline <baseline-id> --json
node dist/index.js gate --run <run-id> --baseline <baseline-id> --json
node dist/index.js gate --run <run-id> --baseline <baseline-id> --wait --timeout-ms 60000 --json
node dist/index.js baseline status --json
node dist/index.js investigation inspect examples/investigations/response-only.json
node dist/index.js investigation inspect examples/investigations/stateful.json --json
node dist/index.js investigation fetch --run <run-id> --evaluation <evaluation-id> --attempt <attempt-id> --criterion <criterion-id> --json
node dist/index.js investigation export-regression examples/investigations/response-only.json --out regression.yamlInspecting an investigation does not quote, execute a target, or consume credits. Reproducing the pinned case is a new quoted run:
node dist/index.js test \
--investigation examples/investigations/response-only.json \
--max-credits 30 \
--yesPublic coverage catalog listing is unauthenticated. Compile and shard
execution are hosted compiler features, not a local pricing engine. Static
catalog counts are informative. POST /v1/billing/quote remains the only
cost preview. --all-shards requires a finite aggregate --max-credits
ceiling and stops before exceeding consent. Interrupted multi-shard work
resumes only with --execution-id against the original API origin, workspace,
and connector. A login switch is SELECTION_TENANT_MISMATCH and does not quote
the next shard. Omitting --execution-id after a terminal attempt starts a new
execution that does not inherit prior shard results. Unbound
aw-selection-execution/2 or spent aw-selection-progress/1 files are not
migrated onto the current login (SELECTION_LEGACY_UNBOUND).
node dist/index.js catalog list --json
node dist/index.js catalog show response-quality/0.1.0/R01
node dist/index.js selection compile \
-c augmentworks.yaml \
--assessment ./augmentworks.assessment.yaml \
--out ./suite-selection.manifest.json
node dist/index.js test \
--manifest ./suite-selection.manifest.json \
--shard shard-000 \
--max-credits 30 \
--yes
node dist/index.js test \
--manifest ./suite-selection.manifest.json \
--all-shards \
--max-credits 30 \
--yes
node dist/index.js test \
--manifest ./suite-selection.manifest.json \
--all-shards \
--execution-id <uuid> \
--max-credits 30 \
--yes
node dist/index.js gate --manifest-file ./suite-selection.manifest.json --declared-shards ./suite-selection.declared-shards.json --jsonSource assessment doctor and quoted hosted execution:
node dist/index.js doctor \
--assessment ./augmentworks.assessment.yaml \
--profile quick
node dist/index.js test \
--assessment ./augmentworks.assessment.yaml \
--profile quick \
--max-credits 30 \
--yes
node dist/index.js test \
--assessment ./augmentworks.assessment.yaml \
--profile full \
--max-credits 30 \
--yesIf grading is pending after target work finishes, evidence is saved. Wait on
the original run; do not re-run the test command. run wait also continues
while target execution is still queued, connected, running, or
cancel_requested, including when grading is absent. run status is not a
full report: it does not page criterion evidence. Export the complete hosted
report with run report <run-id> --json (JSON stdout only;
aw-run-report-export/1). Exit 0 means the assessment passed with complete
required grading and known coverage. A successful status query of unfinished
work is not a pass (ok: true with assessment: "incomplete" and exit 11).
A completed run with a null outcome never exits 0. An expected failing
negative-control report exits 10 and still includes mapped responses and
criterion documents. A new required semantic regression also exits 10 even
when aggregate pass rates match and the comparison HTTP request succeeded.
Pending judging, evaluator error, incompatible scope, and missing required
coverage cannot exit 0. Billing rejection is exit 13 and is not a chatbot
assertion failure. Account-free demo, test --local, offline doctor, and
schema still make no billing calls. doctor is offline validation only; it
does not check AugmentWorks authentication or fetch a hosted report.
compare, gate, and baseline status consume the hosted
aw-release-policy/1 decision. They do not create a run, reserve credits,
call the target, or promote a pin. gate --wait re-queries the original run
ID through run wait semantics, then evaluates that same ID. Promotion is an
explicit baseline promote --expected-revision and is never automatic.
Machine credentials typically cannot promote. The complete hosted GitHub
Actions recipe is docs/examples/github-actions-hosted.yml (scoped
AUGMENTWORKS_API_KEY, isolated synthetic target, --headless, quoted
--max-credits, run wait on the original ID, then gate). A shorter
gate-only snippet remains in docs/examples/github-actions-hosted-gate.yml.
Copy-pastable noninteractive CI (this 0.3.7 package after npm ci && npm run build;
no browser; npx --yes is not a spending ceiling). Capture the run id, wait
on that exact run if grading is pending, and recover an interrupted create
before considering a new admission:
set +e
json=$(node dist/index.js test \
--assessment ./augmentworks.assessment.yaml \
--max-credits 30 \
--yes \
--json)
code=$?
set -e
run_id=$(printf '%s\n' "$json" | node -e "
let s = '';
process.stdin.on('data', (d) => { s += d; });
process.stdin.on('end', () => {
const parsed = JSON.parse(s);
process.stdout.write(String(parsed.run_id ?? parsed.runId ?? ''));
});
")
if [ -z "$run_id" ]; then
echo "No run id. Inspect recover before starting another hosted test." >&2
set +e
node dist/index.js recover --json
set -e
exit "$code"
fi
if [ "$code" -eq 11 ]; then
node dist/index.js run wait "$run_id" --json --timeout-ms 60000
code=$?
fi
# Read-only complete export. Do not run logout in automation cleanup;
# logout revokes reusable workspace API keys.
node dist/index.js run report "$run_id" --json
# Read-only semantic gate. Does not start, reserve, or charge a run.
# Set AUGMENTWORKS_BASELINE_ID to an explicit pin; do not auto-promote.
if [ -n "${AUGMENTWORKS_BASELINE_ID:-}" ]; then
node dist/index.js gate \
--run "$run_id" \
--baseline "$AUGMENTWORKS_BASELINE_ID" \
--json
code=$?
fi
exit "$code"Do not use @latest or immutable npm 0.3.6 for this package. Website
examples adopt independently inspected @augmentworks/cli@0.3.7 (AUG-250).
Independent registry evidence lives in
docs/feature-readiness/published-registry-evidence.json.
No AugmentWorks account, login, credit, relay, or dashboard is required. Point
the published CLI at a target you are authorized to call, or clone this
repository for the fictional refund-agent example server. The bundled starter
packet is synthetic. examples/ is not in the npm tarball. The local CLI
itself is this 0.3.7 package. This path is not the packaged demo command.
git clone https://github.com/jeffskafi/augmentworks-cli.git
cd augmentworks-cli
npm ci
npm run build
cd examples/refund-agent
[ -f .env ] || cp .env.example .env
# Windows: if not exist .env copy .env.example .env
node --env-file=.env server.mjsIn another terminal, from the example directory, run the published local CLI:
npx --yes @augmentworks/cli@0.3.7 doctor \
-c augmentworks.yaml
npx --yes @augmentworks/cli@0.3.7 test \
--local \
-c augmentworks.yaml \
--packet support-refunds-starter@0.1.0 \
--opendoctor still does not contact the target. test --local then calls only the
target selected by the local configuration and writes private JSON, JUnit, and
static HTML reports beneath .augmentworks/runs/<run_id>/. --open opens that
static HTML file. Local reports do not upload to AugmentWorks. Local
test --local requires no AugmentWorks account and contacts no AugmentWorks
service.
“Local” describes the AugmentWorks boundary, not an air gap. The CLI makes no AugmentWorks control-plane request, but the configured target may itself be a network service and may call models or other dependencies.
Interactive credentials use the macOS login Keychain, Windows CurrentUser
DPAPI, or the Linux Secret Service when available. On POSIX systems only, an
explicit --allow-file-credentials opt-in enables a warned mode-0600 local
file when no native store is available. Plaintext fallback is disabled on
Windows because POSIX file modes cannot establish a safe Windows ACL.
Do not put an AugmentWorks token on the command line. Long-lived project-token
issuance is not part of the interactive connector-auth release, so do not
substitute its one-hour interactive access token for an unattended CI
credential. This 0.3.7 package accepts AUGMENTWORKS_API_KEY for noninteractive
workspace keys issued at https://augmentworks.ai/portal/settings/api-keys.
That mode never loads a keychain, refreshes, or persists credentials. Differing
nonempty AUGMENTWORKS_API_KEY and AUGMENTWORKS_TOKEN values fail closed
before any network call. AUGMENTWORKS_TOKEN remains available for paired
token/refresh injection when the API key is unset. CHATBOT_API_KEY is only
the target secret named by YAML bearer_env; it is not a platform
key. Do not run logout from routine CI cleanup.
The CLI loads .env from the directory containing the selected config. YAML
contains environment-variable names, never credential values.
version: 1
target:
name: refunds-staging
connector: http
base_url: ${CHATBOT_BASE_URL}
auth:
bearer_env: CHATBOT_API_KEY
operations:
prepare:
method: POST
path: /__augmentworks/prepare
idempotent: true
request:
attempt_id: $input.attempt_id
fixture: $input.fixture
response:
status: $.status
send:
method: POST
path: /chat
idempotent: false
request:
message: $input.message.content
turn_id: $input.turn_id
attempt_id: $input.attempt_id
response:
content: $.answer
tool_events: $.events
finished: $.finished
observe:
method: POST
path: /__augmentworks/observe
idempotent: true
request:
attempt_id: $input.attempt_id
probe_keys: $input.probe_keys
response:
order.status: $.order.status
order.refunded_amount: $.order.refunded_amount
order.refundable: $.order.refundable
cleanup:
method: POST
path: /__augmentworks/cleanup
idempotent: true
request:
attempt_id: $input.attempt_id
telemetry:
allow_tool_events: true
allow_observations:
- order.status
- order.refunded_amount
- order.refundable# .env
CHATBOT_BASE_URL=http://127.0.0.1:8000
CHATBOT_API_KEY=replace-locallySee the configuration reference, the versioned
augmentworks.yaml schema, and the
refund-agent example.
| Mode | AugmentWorks contact | Evidence boundary |
|---|---|---|
Local (test --local) |
None | Packet inputs, mapped responses, tool events, observations, scoring, and reports stay in the customer environment. Only the locally configured target is contacted. |
Hosted (test) |
Outbound HTTPS authentication and relay | Raw target boundary and secrets stay local. Bounded packet inputs and mapped, allowlisted evidence are exchanged with AugmentWorks. |
The connector credential created by login authenticates the CLI to
AugmentWorks. Target authentication is separate: YAML names local environment
variables whose values are used only for CLI-to-target requests.
sequenceDiagram
participant CLI as Customer-run CLI
participant Cloud as AugmentWorks relay
participant App as Test application
CLI->>Cloud: Create run and long-poll over HTTPS
Cloud-->>CLI: One typed operation
CLI->>App: Locally configured HTTP request
App-->>CLI: Application response
CLI->>Cloud: Mapped, bounded result
- The CLI authenticates and creates an assessment run.
- It long-polls the relay over outbound HTTPS; no inbound port or tunnel is required.
- The relay can request only
prepare,send,observe, orcleanup. - Local configuration—not a cloud command—selects the URL, method, headers, environment variables, and response mapping.
- The CLI returns only mapped assistant content, opted-in tool events, allowlisted observations, safe errors, and lifecycle status.
- AugmentWorks evaluates that evidence and updates the dashboard.
- When a fixture may exist, the hosted relay can dispatch a typed
cleanupfollow-up after success, failure, cancellation, or interruption. The CLI does not invent lifecycle operations that the relay did not dispatch.
The relay delivers commands at least once. The CLI journals command IDs and returns a durable terminal result for an exact duplicate. It never blindly retries an ambiguous operation unless local configuration explicitly declares that operation idempotent.
Before run creation, test validates the config locally, resolves the current
workspace and connector identity, and persists a secret-free active intent.
Re-running the exact command on the same machine resumes the same run. A
different active packet, configuration, connector, or workspace is refused
instead of silently creating another run or reserving another credit.
- The CLI validates the YAML, resolves target credentials locally, and loads
either the bundled
support-refunds-starter@0.1.0packet or a localpacket.json. - It verifies that the configured lifecycle mappings and telemetry allowlist satisfy the packet's declared capabilities.
- It executes attempts serially:
prepare, one or moresendoperations,observe, andcleanupin afinallypath. - It deterministically evaluates the packet assertions and writes
report.json,junit.xml, andreport.htmlto a fresh private directory. - If
--openis present, it opens the generated static HTML report. No dashboard or hosted evidence record is created.
The CLI does not blindly retry an ambiguous non-idempotent operation. It still attempts observation when a send outcome is ambiguous and attempts cleanup when a fixture may exist. A cleanup failure stops new attempts. The first Ctrl+C requests cancellation and drains cleanup; a second exits immediately. A hard process or machine failure cannot guarantee cleanup, so synthetic fixtures need an independent server-side TTL. New target work is bounded by a 30-minute local run deadline; bounded cleanup is still allowed to drain after that deadline.
Local packets use the strict aw-packet/0.1 JSON format. They are data, not
executable plugins: no JavaScript, modules, shell commands, remote URLs, or
download step is accepted. --packet may name the bundled
support-refunds-starter@0.1.0, a local JSON file, or a local directory
containing packet.json. Packets with schema_version: "aw-packet/0.2",
evaluation_mode: hybrid, or llm_rubric criteria are refused in --local
before any target call.
Print the packet and result schemas with:
npx --yes @augmentworks/cli@0.3.7 schema --kind local-packet
npx --yes @augmentworks/cli@0.3.7 schema --kind local-resultThe default exact output directory is .augmentworks/runs/<run_id>. Override
it with --output-dir <path>; the selected leaf must not already exist, and the
CLI never merges into or overwrites an existing directory. On POSIX systems the
leaf is mode 0700 and report files are mode 0600.
Every local report is labeled:
Local, customer-executed result. AugmentWorks did not receive or independently verify this run. This artifact is unsigned and is not a certification, audit, or hosted evidence record.
The JSON uses AW-LOCAL-RESULT-1 and includes a SHA-256 change-detection
checksum. That checksum is not a signature or proof of provenance. The HTML is
a self-contained static file with no scripts or external assets. Treat all
three artifacts as sensitive customer-controlled evidence. --json emits the
same final local result on stdout; the three files are still generated.
--assessment is in this @augmentworks/cli@0.3.7 package. init writes
augmentworks.yaml, augmentworks.assessment.yaml, and referenced starter
files before you run an assessment. See examples/response-agent/ for the FAQ
fixture used as the packaged response-quality starter.
npx --yes @augmentworks/cli@0.3.7 doctor \
--assessment ./augmentworks.assessment.yaml \
--profile quick
npx --yes @augmentworks/cli@0.3.7 test \
--assessment ./augmentworks.assessment.yaml \
--profile quick \
--open
npx --yes @augmentworks/cli@0.3.7 test \
--assessment ./augmentworks.assessment.yaml \
--profile full \
--openThe assessment file is aw-assessment-file/1 YAML. It selects packet versions
and optional scenario IDs, a profile (quick, full, combined, or custom),
and evaluation_mode (deterministic or hybrid). It may attach local
.md/.txt reference files under the assessment directory or hosted reference
ids, not both. It never contains judging credentials. Hosted --assessment
uses relay protocol aw-relay/0.2, freezes the YAML and reference bytes, and
refuses resume if those hashes change. --assessment cannot be combined with
--local. If grading is still pending after target work, the CLI exits 11
and does not print a hybrid pass.
See examples/response-agent/ for a synthetic FAQ assessment file.
| Command | Purpose | Side effects |
|---|---|---|
login [--device] [--workspace uuid] [--allow-file-credentials] |
Authorize this machine | Opens a browser by default and stores a revocable credential. --workspace must match /auth/me before the origin slot is replaced |
logout |
Revoke and remove the connector credential | Requests server-side revocation and deletes local credential material. Do not use as routine CI cleanup when AUGMENTWORKS_API_KEY is set |
whoami [--workspace uuid] |
Show the current workspace identity | Reads cloud identity; may refresh interactive credentials. Prints name and UUID. API-key mode reports principal kind, credential id, actions, and expiry without the bearer |
usage [--workspace uuid] [--json] |
Show authenticated workspace execution-credit usage | Read-only billing snapshot; no target YAML, grant, reservation, checkout, subscribe, or cancel |
billing [--json] [--print] [--open] |
Open or print the first-party billing page | Read-only navigation; no Checkout, Stripe customer, refund, subscription mutation, or reservation |
init [-c path] [--starter name] [--agent] [--force] |
Generate config, assessment, starter references, and setup guidance | Writes the complete starter. Does not overwrite edited files unless --force is explicit; never replaces .env |
doctor [-c path] [--offline] [--json] [--assessment path] [--profile profile] |
Validate config, mappings, secrets, local prerequisites, assessment files, and wire bounds | Makes no network calls, invokes no lifecycle hook, and consumes no assessment credit |
preview-mapping [-c path] [--operation kind] [--fixture path] [--probe-keys keys] [--json] |
Preview response mappings and the exact sanitized evidence payload from a local JSON fixture | Reads only the selected config and fixture. No target, cloud, or model call |
probe [-c path] [--yes] [--json] |
Explicit bounded connection probe | Prints the planned calls first. --yes executes them against the configured target only and consumes real target allowance. Never runs during doctor or init. No hosted API, no credits |
catalog list [--json] / catalog show <id> [--json] |
List or show the public coverage catalog | Unauthenticated. Static counts are informative, not a quote. Stale --catalog-version fails closed |
selection compile [--assessment path] [--profile smoke|release] [--json] [--out path] |
Ask the server compiler for included/excluded/incompatible cases and bounded shards | Not a quote and not a billable run. Cannot be used with --local |
suite validate <file> [--json] / suite preview <file> [--json] / suite preflight <file> [-c path] [--json] |
Validate, preview, or preflight a customer-owned aw-suite/1, aw-suite/2, or aw-suite/3 file |
Offline; not a price; does not execute a target, buy credits, or call an LLM. Live origin matching is optional via -c. Hosted aw-suite/3 quote stays fail-closed until the real-data release gate. See docs/customer-suites.md |
test [-c path] --packet name@version [--open] |
Run one hosted assessment | Authenticates to AugmentWorks, calls configured lifecycle endpoints, and may create synthetic state |
test [-c path] --assessment path [--profile profile] [--workspace uuid] [--estimate] [--max-credits n] [--yes] [--open] |
Quote or run a hosted assessment from an assessment file | Uses aw-relay/0.3 quotes. --estimate never reserves credits. npx --yes is not a spending ceiling. Optional YAML selection (smoke/release) is compiled by the server. --workspace is compared to /auth/me before quote |
test [-c path] --manifest path [--shard id | --all-shards] [--execution-id uuid] [--max-credits n] [--yes] [--artifact-out path] |
Quote or run one compiled shard, or a finite multi-shard driver | Prints exclusions before consent. --all-shards requires --max-credits. Resume an unfinished attempt with --execution-id. After it is terminal, omit --execution-id to start a new rerun. Cannot be used with --local |
test [-c path] --suite path [--estimate] [--max-credits n] [--yes] [--headless] [--open] |
Quote or run a hosted customer-owned suite | Pins the server-accepted revision. Changing the file after quote does not silently alter admitted work. --suite cannot be used with --local. --headless requires AUGMENTWORKS_API_KEY or AUGMENTWORKS_TOKEN and never opens a browser |
investigation inspect <file> [--json] / investigation fetch --run id --evaluation id --attempt id --criterion id [--out path] [--json] / investigation export-regression <file> --out path [--json] |
Inspect or download a safe failure investigation, or export a reviewed aw-suite/1 regression draft |
Observation only. Does not execute a target, shell fragment, evaluator, or quote. Copied commands are data. See docs/investigation.md |
test [-c path] --investigation path [--estimate] [--max-credits n] [--yes] [--open] |
Reproduce the exact pinned case from an investigation file | New quote and consent every time. Never selects latest or reuses a consumed quote. Cannot be combined with --local, --suite, --assessment, or --packet |
run status <run-id> / run wait <run-id> / run retry-evaluation <run-id> / run report <run-id> [--scope live-informational] |
Inspect, wait, retry incomplete grading, or export the complete hosted report | Status/wait/report are read-only. run report writes one aw-run-report-export/1 JSON document unless --scope live-informational requests the separate aw-run-report-live-scope/1 overlay. Retry-evaluation debits 0 customer credits and does not replay the target |
compare --run <run-id> --baseline <baseline-id> [--json] |
Compare a candidate run against an explicit pinned baseline | Read-only. Does not start a test, reserve credits, or consume credits. Rejects missing identities; does not invent a pin |
gate --run <run-id> --baseline <baseline-id> [--wait] [--timeout-ms n] [--json] |
Evaluate the hosted aw-release-policy/1 release decision |
Transport success is not a pass. --wait re-queries the original run ID only. A finalized machine-principal gate is the CI provenance record. See docs/examples/github-actions-hosted.yml |
gate --manifest-file path --declared-shards path [--wait] [--json] |
Evaluate the hosted whole-suite aw-manifest-release-policy/2 receipt |
Exit 0 only for a server-authoritative exact-coverage pass. Empty, non-executable, or incomplete declarations never reach the network. Legacy v1 receipts fail closed. |
baseline status [--json] / baseline promote --run <run-id> --baseline <id> --expected-revision <n> [--json] |
List pins, or explicitly promote a candidate onto a pin | Status is read-only. Promote is never automatic, requires --expected-revision, and stays off the machine allowlist unless the server grants baseline:promote |
recover [-c path] [--retire | --resume | --cancel] [--json] |
Inspect or recover a hosted assessment | Does not create a new run. Default inspection only; --retire, --resume, and --cancel are mutually exclusive. Do not delete journals when admission is unknown |
action recover --state-dir path [--origin url] [--json] |
Retry a durable controlled-action receipt without rerunning the customer tool | Public controlled actions stay unavailable. Offline mode makes no hosted call and claims no platform signature. A rejected receipt stays pending. See docs/action-gate.md |
demo [--json] [--open] [--output-dir path] [--mode full|faulty|corrected] |
Packaged loopback refund demonstration | Contacts only an isolated 127.0.0.1 target owned by this command |
test --local [-c path] --packet reference [--output-dir path] [--open] [--json] |
Run and score a customer-executed local assessment | Contacts only the configured target and writes local artifacts; no AugmentWorks account or service is used. --workspace is rejected; AUGMENTWORKS_WORKSPACE_ID is ignored |
schema [--kind config|local-packet|local-result|customer-suite|investigation-export] |
Print a bundled v1 JSON Schema | None |
| Code | Meaning |
|---|---|
0 |
Assessment passed, or a compatible release gate returned policy pass. For demo (default --mode full), the fail-then-pass story and cleanup succeeded. |
1 |
Internal or report-generation failure, or a demo whose faulty run unexpectedly passed |
2 |
Configuration, packet, capability, output preflight, incompatible comparison scope, or a missing required baseline pin |
3 |
Hosted authentication failure; unreachable from --local |
4 |
Hosted relay/protocol failure, including an authorized baseline promotion conflict; unreachable from --local |
5 |
Target, protocol-evidence, or indeterminate execution error |
6 |
Cleanup failure; takes precedence over assessment status |
10 |
Assertions failed, the assessment was inconclusive, or the release policy blocked (including a new required semantic regression while aggregate pass rates match) |
11 |
Hosted grading is pending, partial, unknown, unsupported, a report export is incomplete, or the release gate is incomplete |
12 |
Required hosted evaluation did not complete because of an operational judging error |
13 |
Hosted billing/usage rejection; unreachable from --local |
130 |
Interrupted after cleanup was drained; a second interrupt exits immediately |
| Level | Required mapping | Evidence claim |
|---|---|---|
| Chat-only | send request and response |
Conversational behavior |
| Tool-aware | send plus structured tool events |
What the chatbot attempted to invoke |
| Stateful | prepare, send, observe, and cleanup |
Values returned by the configured synthetic-state observer |
For consequential workflows, an agent saying “done” is not proof. Stateful evidence requires configured observation and cleanup hooks. A hosted run records what the customer-controlled observer reports in AugmentWorks; a local run records it only in the customer-held reports. Neither mode independently proves that the observer is truthful or that staging matches production.
- The CLI accepts only typed lifecycle operations. It does not accept shell commands, file instructions, arbitrary URLs, methods, headers, or modules from the relay.
- Secrets are resolved locally and redacted from diagnostics. Interactive credentials use the operating-system credential store when available.
- Telemetry is opt-in and allowlisted. Run preflight sends sorted public observation aliases—not values, selectors, environment-variable names, or target URLs. Request and response sizes, timeouts, and nesting are bounded.
- Run preflight sends an unkeyed checksum of the resolved target boundary, not its raw URL or paths. It detects boundary drift across local restart; it is not target identity, ownership, code, state, or execution proof.
- Executed prompts and synthetic fixtures are visible to the local connector; undispatched packet branches and hosted assertions can remain private.
- A local packet is fully visible to the customer process, and its unsigned report is customer-controlled. It must not be presented as hosted or independently verified AugmentWorks evidence.
- The CLI's unkeyed SHA-256 digests detect a conflicting replay when compared with an already-durable local or relay record; they are not evidence signatures. The hosted relay associates accepted results with authenticated connector, session, run, packet, configuration, and sequence bindings.
- A customer-operated observation hook can be incorrect or dishonest, and a staging result is not proof of production equivalence.
- v0.2 hosted runs in this package use an authorized, isolated synthetic or staging target and constructed test data while hosted real-data is release-disabled. See Data scope.
Read the complete security model, relay protocol, and security policy.
Test your chatbot with your questions and business rules. That hosted scope
stays off until a compatible published release is enabled. Until then, hosted
assessments use an authorized,
isolated synthetic or staging target and constructed test data. Hosted
real-data quote and admission stay release-disabled (EXECUTION_RELEASE_UNAVAILABLE).
Do not point a hosted run at a production system, and do not upload production,
customer, or regulated records, while that release is disabled.
When the release is enabled, a customer administrator authorizes an exact owned
target (owned_target or written_permission) and a data policy for public,
business, or ordinary personal content. Hostname control is not ownership.
Secrets, PHI, payment-card credentials, and other separately restricted data
stay out of evaluation content. This text is not legal approval and does not
claim SOC 2, HIPAA, Zero Data Retention, or guaranteed reversal of actions.
Offline aw-packet/0.1 stays synthetic-only and account-free.
aw-packet/local-authorized-1 is a customer-declared local scope that never
contacts AugmentWorks and is not hosted authority. Older synthetic and
live-informational files still load. Packaged examples and starters remain
fictional. The minimum published npm artifact for a later real-data release is
owned by the publication gate; this package does not guess that version.
After the packaged demo, configure the generic HTTP connector. While hosted real-data is release-disabled, use an authorized, isolated synthetic or staging target:
npx --yes @augmentworks/cli@0.3.7 init --agentnode dist/index.js initornode dist/index.js init --starter workflowfrom a clone afternpm ci && npm run build. Map only the hooks required by that pattern.- Put secret names in YAML and values only in local
.env. npx --yes @augmentworks/cli@0.3.7 doctor -c augmentworks.yamlnode dist/index.js preview-mapping -c augmentworks.yaml --operation send --fixture ./fixtures/send-response.jsonnode dist/index.js probe -c augmentworks.yamlthennode dist/index.js probe -c augmentworks.yaml --yesnpx --yes @augmentworks/cli@0.3.7 test --local -c augmentworks.yaml --packet support-refunds-starter@0.1.0
Do not fabricate an OpenAI, LangServe, MCP, or framework adapter the CLI does not provide. Hosted access remains an invited workspace at https://augmentworks.ai; this package does not define pricing or public signup.
- The generic HTTP connector is the only v0.2 connector.
- OpenAPI import, OpenAI-compatible presets, LangServe, and custom modules are not implemented.
- v0.2 exposes no public
connectcommand; hostedtestkeeps the connector online only for the assessment it starts. - Re-running the same hosted
testcommand resumes an active bound intent when admission already succeeded. If payment or create state is unknown, inspect withrecoverand wait on the original run; do not delete journals or blindly rerun the test. There is no--rerunflag. - Pointing the CLI directly at a model provider tests the model endpoint, not the customer's policies, tools, database, or application behavior.
- This
@augmentworks/cli@0.3.7package includes--assessment,demo, andinitstarter generation. Packagedusage,billing,preview-mapping,probe,test --estimate,--max-credits,run status/run wait/run report,compare/gate/baseline,AUGMENTWORKS_API_KEYmode, owned-suite commands, investigation inspect/fetch, and own-target starters are in this package. Immutable npm0.3.4remains published; do not overwrite it. Immutable npm0.3.3is not this release line.
npm ci
npm run typecheck
npm test
npm run build
npm run smoke:packSee CONTRIBUTING.md for repository conventions and the agent setup guide for a safe coding-assistant workflow.
The package requires Node.js 20+. Configuration and relay envelopes carry an explicit protocol version. Patch releases remain compatible with their v1 schema; incompatible configuration or protocol changes require a new version. Published releases are expected to use npm trusted publishing with provenance.
See CHANGELOG.md for release notes.
Apache-2.0. See LICENSE.