Tier: L (5-7 days) | Type: docs
Context. guides/ops/self-hosted-deployment.mdx (Wave 7) walks getting contracts up. Nothing today tells the operator what to watch, when to page, or how to respond. Teams running Wraith in production need a full runbook: healthy metric baselines, ready-to-import dashboards, alert rules, incident playbooks, and a tabletop exercise. This is more than a page; it is a bundle of ops artifacts, hence the larger tier.
Scope.
- New
guides/ops/monitoring-and-on-call.mdx.
- Key metrics: RPC latency, announcement-lag, indexer backlog, watcher event-drop rate, contract error rate.
- Grafana dashboard JSON in
guides/ops/dashboards/ (one for RPC, one for indexer, one for contracts).
- Prometheus alerting rules file with rationale per rule.
- Six named incident playbooks: RPC outage, Horizon backpressure, indexer stall, contract mispublish, watcher drop-spike, key-rotation incident.
- Sev matrix mapping incident types to response SLA, cross-linked with the auditor guide's severity ladder.
- Tabletop exercise script for a quarterly drill.
Acceptance.
Files. guides/ops/monitoring-and-on-call.mdx (new), guides/ops/dashboards/ (new), guides/ops/alerts/ (new), docs.json.
Tier: L (5-7 days) | Type: docs
Context.
guides/ops/self-hosted-deployment.mdx(Wave 7) walks getting contracts up. Nothing today tells the operator what to watch, when to page, or how to respond. Teams running Wraith in production need a full runbook: healthy metric baselines, ready-to-import dashboards, alert rules, incident playbooks, and a tabletop exercise. This is more than a page; it is a bundle of ops artifacts, hence the larger tier.Scope.
guides/ops/monitoring-and-on-call.mdx.guides/ops/dashboards/(one for RPC, one for indexer, one for contracts).Acceptance.
promtool check rulesmint devCompile docs snippetsCIself-hosted-deployment.mdxandreference/auditor-guide.mdxFiles.
guides/ops/monitoring-and-on-call.mdx(new),guides/ops/dashboards/(new),guides/ops/alerts/(new),docs.json.