From 54eb02cd780c5b34fb8bf3e691ca937cb7f3f668 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Mon, 28 Sep 2026 23:23:38 +0100 Subject: [PATCH 01/25] docs(clients): document QWP ingestion and queries in the JavaScript client Rewrite the Node.js client page as the JavaScript client page for @questdb/nodejs-client 5.0.0 and the new @questdb/browser-client package: pooled ingestion and streaming SQL queries, column types, compiled writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, browser session authentication, and migration from ILP and 4.x. Mark Node.js QWP support as beta on the Connect overview and client cards, and note where the JavaScript client differs on the connect-string, store-and-forward, failover, and client-behavior pages. Document the qwp.browser.tls.termination.enabled server setting. --- documentation/changelog.mdx | 2 + documentation/configuration/qwp.md | 27 +- .../connect/clients/connect-string.md | 47 +- .../clients/date-to-timestamp-conversion.md | 47 +- documentation/connect/clients/nodejs.md | 2661 ++++++++++++++++- .../connect/compatibility/pgwire/nodejs.md | 13 +- documentation/connect/overview.md | 34 +- .../wire-protocols/qwp-client-behavior.md | 10 +- .../wire-protocols/qwp-egress-websocket.md | 4 +- .../wire-protocols/qwp-ingress-websocket.md | 4 +- .../client-failover/configuration.md | 4 +- .../store-and-forward/concepts.md | 2 +- .../store-and-forward/configuration.md | 4 +- documentation/sidebars.js | 2 +- shared/clients.json | 12 +- 15 files changed, 2638 insertions(+), 235 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 833739bcce..ddc3db65f7 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -27,9 +27,11 @@ This page tracks significant updates to the QuestDB documentation. ### Reference - Added [`cairo.sql.subsample.max.rows`](/docs/configuration/cairo-engine/#cairosqlsubsamplemaxrows), the input row limit for the `lttb`, `m4`, `minmax`, `uniform`, and `cadence` methods of `SUBSAMPLE`. It does not apply to `sdt` +- Added [`qwp.browser.tls.termination.enabled`](/docs/configuration/qwp/#qwpbrowsertlsterminationenabled), for browser QWP connections behind a TLS-terminating reverse proxy, and documented the same-origin check QuestDB applies to [browser connections](/docs/configuration/qwp/#browser-connections) ### Updated +- [JavaScript client](/docs/connect/clients/nodejs/) - Rewrote the Node.js client page for QWP support in `@questdb/nodejs-client` 5.0.0 and the new `@questdb/browser-client` package: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, browser session authentication, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) now lists where the JavaScript client's keys and defaults differ - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/configuration/qwp.md b/documentation/configuration/qwp.md index 1928bcf1d1..03d72df52b 100644 --- a/documentation/configuration/qwp.md +++ b/documentation/configuration/qwp.md @@ -6,7 +6,8 @@ description: QWP is QuestDB's columnar binary protocol for high-throughput data ingestion (`/write/v4`) and streaming query results (`/read/v1`) over WebSocket and UDP. -These properties control protocol limits and the UDP receiver. WebSocket +These properties control protocol limits, query result compression, browser +connections, and the UDP receiver. WebSocket ingestion and egress share the HTTP server's network settings (port, TLS, worker threads); see [HTTP server configuration](/docs/configuration/http-server/) for those. @@ -52,6 +53,30 @@ overriding the level the client requests via `X-QWP-Accept-Encoding`. `0` disables the override and honours the client's request. Any other value must be in the range `1`-`9`; the server refuses to start otherwise. +## Browser connections + +Browser applications, such as those using the +[JavaScript client](/docs/connect/clients/nodejs/#browser-applications), open +QWP WebSockets from a web page. Browsers always send an `Origin` header with +the upgrade, and QuestDB accepts it only when it is same-origin with the +request's `Host`, including the scheme: an `http://` origin is refused over TLS, +and an `https://` origin is refused over plain HTTP. This blocks cross-site +WebSocket hijacking. Upgrades without an `Origin` header, which is how +non-browser clients connect, are unaffected. Browser QWP connections require a +QuestDB release newer than 10.0.1. + +### qwp.browser.tls.termination.enabled + +- **Default**: `false` +- **Reloadable**: no + +Treats plain-HTTP connections as secure for the browser origin check. Enable it +when a reverse proxy terminates TLS in front of QuestDB: an `https://` origin is +then accepted, and an `http://` origin is refused. The proxy must forward the +browser's original `Host` header unchanged, including the port. The setting +affects the origin check only; it does not mark the `qdb_session` cookie as +`Secure`. + ## UDP receiver :::note diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index c027700dd5..0bd62b6b08 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -256,6 +256,7 @@ Selecting the `wss` schema enables TLS. | Rust, C, C++, Python | PEM (default), JKS, PKCS#12 | | .NET | PKCS#12 / PFX | | Go | none — OS trust store only, both keys rejected at parse time | + | JavaScript (Node.js) | PEM only; `tls_roots_password` is rejected at parse time | - `tls_roots_password` — password for the `tls_roots` file. Required only for a JKS or PKCS#12 trust store; PEM needs no password. Setting it without @@ -291,7 +292,9 @@ Existing JKS and PKCS#12 trust stores keep working through The Go client verifies against the operating-system trust store only and **rejects both keys at parse time**; to trust a private CA there, install it in -the host trust store. On Rust, C, C++ and Python, `tls_roots_password` switches +the host trust store. The JavaScript client on Node.js accepts PEM only and +rejects `tls_roots_password`: export a JKS or PKCS#12 trust store to PEM +first. On Rust, C, C++ and Python, `tls_roots_password` switches the file to a Java keystore and is QWP/WebSocket only: other transports keep PEM as the sole format. Check the relevant [client library page](/docs/connect/overview/#client-libraries) for @@ -320,15 +323,17 @@ act independently: whichever threshold trips first sends the batch. buffered rows. - `auto_flush_rows` — flush when the buffered row count reaches this threshold. Set to `off` to disable. Default where supported: `1000`. + The JavaScript client rejects `off` here; set `0` to disable. - `auto_flush_interval` — flush when this many milliseconds have elapsed since the first buffered row. The client evaluates the interval on the next `at()` / `flush()` call, not on a wall-clock timer. Set to `off` to - disable. Default where supported: `100` (100 ms). + disable. Default where supported: `100` (100 ms). The JavaScript client + rejects `off` here; set `0` to disable. - `auto_flush_bytes` — flush when the encode buffer reaches this byte size. Set to `off` to disable. Accepts [size suffixes](#size-suffixes). **The default differs by client**: Java - ships it **disabled** (`0`), .NET defaults to `8m` (8 MiB), and Rust, C and - C++ reject the key outright. A Java application that assumes an 8 MiB byte + and JavaScript ship it **disabled** (`0`), .NET defaults to `8m` (8 MiB), and + Rust, C and C++ reject the key outright. A Java application that assumes an 8 MiB byte trigger is active will size batches expecting a flush that never fires. When set to a positive value, the client clamps the effective threshold down to 90% of the server- @@ -381,6 +386,9 @@ case-insensitive and 1024-based, matching `-Xmx` conventions: | `g` or `gb` | GiB (× 1024³) | `1g`, `10gb` | | `t` or `tb` | TiB (× 1024⁴) | `1t` | +The JavaScript client accepts only the single-letter suffixes `k`, `m`, `g`, +and `t`, and rejects `kb`, `mb`, `gb`, and `tb`. + ## Multi-host failover {#failover-keys} *Applies to: ingress and egress. The [Role filter and zone preference](#role-filter-and-zone-preference) @@ -423,6 +431,16 @@ single-primary cluster: ingress automatically follows the primary across the host list and adapts when the primary moves to another node. Ingress silently accepts these keys and ignores them. +:::caution JavaScript client + +The JavaScript client applies `target` and `zone` to ingress as well. With +`target=replica` in a shared connect string, its senders accept only replicas +and cannot ingest. Set the query-side role through the typed `egress` option +instead; see the +[JavaScript client page](/docs/connect/clients/nodejs/#multiple-endpoints). + +::: + - `target` — server-role filter applied per endpoint after the upgrade reads `SERVER_INFO`. Options: - `any` (default) — no preference; route to any healthy endpoint. @@ -500,6 +518,7 @@ equivalent — same architecture, no durability across restarts. |---|---| | Java (`QuestDB` facade) | `/-/` | | Rust, C, C++ (`QuestDb` / `questdb::pool` / `questdb_db`) | `/-ingest-/` | + | JavaScript (`connectQwpNodeClient`) | `/-/` | The minted names belong to that pool's namespace, so pools sharing one `sf_dir` need distinct bases; the slot-in-use error covers both cases @@ -514,7 +533,8 @@ equivalent — same architecture, no durability across restarts. is lost on power failure. The .NET client is the exception: it accepts only `memory` and rejects - anything else at parse time. + anything else at parse time. The JavaScript client also accepts `append`, + which makes every journal append durable before `flush()` resolves. - `sf_sync_interval_millis` — cadence at which `sf_durability=periodic` checkpoints published frames to stable storage. Default: `5000`. Requires `sf_durability=periodic`; rejected otherwise. The configured interval is a @@ -616,6 +636,9 @@ architecture is that a producer survives an arbitrarily long outage. constructor gives up and returns the error. The running loop and the `async` initial connect never consult it. Default: `300000` (5 min). Setting this enables `initial_connect_retry=on` implicitly; see below. + The JavaScript client differs: a sender with neither `sf_dir` nor + `initial_connect_retry=async` applies this budget to every outage, and fails + with `QwpReconnectExhaustedError` when it runs out. - `initial_connect_retry` — whether the client retries the initial connect attempt on failure. - `off` (default, alias `false`) — fail fast on initial connect failure. @@ -636,7 +659,7 @@ architecture is that a producer survives an arbitrarily long outage. milliseconds waiting for buffered frames to drain. Set to `0` or `-1` for fast close (skip the drain). **The default differs by client**: `60000` (60 s) on Java and .NET, `5000` (5 s) on Rust, C, C++ and Python, which - share the same Rust core. + share the same Rust core, and on JavaScript. This is the shutdown data-loss window. Setting it to `0` skips the drain entirely and drops un-ACKed batches on every clean shutdown. @@ -748,7 +771,7 @@ per-language names. *Applies to: the pooled facade (`QuestDB.connect`, `questdb::pool`, `QuestDb::connect`, `questdb.connect`, `qdb.NewQuestDB`, -`QuestDBClient.Connect`).* +`QuestDBClient.Connect`, `connectQwpNodeClient`).* Every client now leads with a pooled facade, so these keys are a first-contact concern. The `Sender` and query-client parsers accept and ignore them; the @@ -791,7 +814,7 @@ consumed by the application. Every client's parser accepts the six `on_*_error` keys below, but only clients that implement the policy layer act on them. **In the Java reference -client they are currently accepted no-ops** — setting +client and the JavaScript client they are currently accepted no-ops** — setting `on_write_error=retriable_other` parses cleanly and changes nothing. .NET does implement them, via `SenderErrorPolicy` and `SenderErrorCategory`. The category table and precedence model below describe the target contract. @@ -854,17 +877,17 @@ description and behaviour notes. | `addr` | `host:port[,host:port…]` | required | [Multi-host failover](#failover-keys) | | `auth_timeout_ms` | int (ms) | `15000` | [Authentication](#auth) | | `auto_flush` | enum (`on` / `off`) | `on` (Rust: only `off`) | [Auto-flushing](#auto-flush) | -| `auto_flush_bytes` | size | Java `0` (off) / .NET `8m` (Rust: rejected) | [Auto-flushing](#auto-flush) | +| `auto_flush_bytes` | size | Java, JavaScript `0` (off) / .NET `8m` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `auto_flush_interval` | int (ms) / `off` | `100` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `auto_flush_rows` | int / `off` | `1000` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `buffer_pool_size` | int (≥ 1) | `4` | [Query client keys](#egress-keys) | | `catch_up_cap_gap_min_escalation_window_millis` | int (ms) | `300000` (5 min) | [Store-and-forward](#sf-keys) | | `client_id` | string | client-specific | [Query client keys](#egress-keys) | -| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python `5000` | [Ingress reconnect](#reconnect-keys) | +| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, JavaScript `5000` | [Ingress reconnect](#reconnect-keys) | | `compression` | enum (`raw` / `zstd` / `auto`) | `raw` | [Query client keys](#egress-keys) | | `compression_level` | int (`1`–`22`) | `1` | [Query client keys](#egress-keys) | | `connect_timeout` | int (ms, `> 0`) | unset | [Authentication](#auth) | -| `connection_listener_inbox_capacity` | int (≥ 1) | `64` (Java) · `256` (Go, .NET) · not supported by Rust, C/C++, Python | [Error handling](#error-handling) | +| `connection_listener_inbox_capacity` | int (≥ 1) | `64` (Java, JavaScript) · `256` (Go, .NET) · not supported by Rust, C/C++, Python | [Error handling](#error-handling) | | `drain_orphans` | enum (`on` / `off`) | `off` | [Store-and-forward](#sf-keys) | | `durable_ack_keepalive_interval_millis` | int (ms) | `200` | [Durable ACK](#durable-ack) | | `error_inbox_capacity` | int (≥ 16) | `256` | [Error handling](#error-handling) | @@ -907,7 +930,7 @@ description and behaviour notes. | `sender_pool_min` | int | `1` | [Connection pool](#pool-keys) | | `sf_append_deadline_millis` | int (ms) | `30000` (30 s) | [Store-and-forward](#sf-keys) | | `sf_dir` | path | unset (memory mode) | [Store-and-forward](#sf-keys) | -| `sf_durability` | enum (`memory` / `periodic`) | `memory` (.NET: `memory` only) | [Store-and-forward](#sf-keys) | +| `sf_durability` | enum (`memory` / `periodic`) | `memory` (.NET: `memory` only; JavaScript also accepts `append`) | [Store-and-forward](#sf-keys) | | `sf_max_segment_bytes` | size | `4 MiB` | [Store-and-forward](#sf-keys) | | `sf_max_total_bytes` | size | `128 MiB` mem / `10 GiB` SF | [Store-and-forward](#sf-keys) | | `sf_sync_interval_millis` | int (ms) | `5000` | [Store-and-forward](#sf-keys) | diff --git a/documentation/connect/clients/date-to-timestamp-conversion.md b/documentation/connect/clients/date-to-timestamp-conversion.md index 9e6efb7286..5e1e2d13fc 100644 --- a/documentation/connect/clients/date-to-timestamp-conversion.md +++ b/documentation/connect/clients/date-to-timestamp-conversion.md @@ -319,28 +319,41 @@ class Program ``` Learn more about the [QuestDB .NET Client](/docs/connect/clients/dotnet/) -## Date to Timestamp in JavasScript/Node.js +## Date to Timestamp in JavaScript/Node.js -The Date type stores both date and time information. - -The QuestDB Node.js client accepts an epoch in microseconds, which can be a `number` or `bigint`. +A JavaScript `Date` stores milliseconds since the Unix epoch. The QuestDB +JavaScript client takes a timestamp as an integer `number` or a `bigint` +together with a unit: `"ms"`, `"us"` (the default), or `"ns"`. A `Date` +therefore needs no arithmetic: pass `getTime()` with the `"ms"` unit. ```javascript -const { Sender } = require("@questdb/nodejs-client") - -const dateStr = '2024-08-05'; -const dateObj = new Date(dateStr + 'T00:00:00Z'); - -// Convert to timestamp (milliseconds since Epoch) then convert to microseconds -const timestamp = BigInt(dateObj.getTime()) * 1000n; -console.log("Date:", dateObj.toISOString().split('T')[0]); -console.log("Timestamp (microseconds):", timestamp.toString()); - -// You can now add the column using QuestDB client, as in -// .timestampColumn("NonDesignatedTimestampColumnName", timestamp) +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const tradeDate = new Date("2024-08-05T00:00:00Z"); + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .timestampColumn("trade_date", tradeDate.getTime(), "ms") + .doubleColumn("price", 2615.54) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} ``` -Learn more about the [QuestDB Node.js Client](/docs/connect/clients/nodejs/) +For an explicit microsecond value, convert through `bigint`: +`BigInt(tradeDate.getTime()) * 1000n`. Nanosecond timestamps, with the `"ns"` +unit, must be a `bigint`. + +Learn more about the [QuestDB JavaScript client](/docs/connect/clients/nodejs/) ## Date to Timestamp in Ruby diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index de60fe5610..ee12a7e2c6 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -1,269 +1,2598 @@ --- slug: /connect/clients/nodejs -title: Node.js Client Documentation +title: JavaScript client for QuestDB +sidebar_label: JavaScript description: - "Get started with QuestDB using the Node.js client for efficient, - high-performance insert operations. Achieve unparalleled time series data - ingestion and query capabilities." + "QuestDB JavaScript client for Node.js and browsers: high-throughput data + ingestion and streaming SQL queries over QWP, with pooling, failover, and + store-and-forward." --- -import { ILPClientsTable } from "@theme/ILPClientsTable" +import SfDedupWarning from "../../partials/_sf-dedup-warning.partial.mdx" + +The QuestDB JavaScript client connects Node.js and browser applications to +QuestDB over [QWP](/docs/connect/wire-protocols/qwp-ingress-websocket/), the +QuestDB Wire Protocol: a columnar binary protocol carried over WebSocket. The +same client ingests data at high throughput and runs SQL queries whose results +stream back as typed, column-oriented batches. + +The client ships as two npm packages built from one code base, with the same +API for ingestion and queries: + +| Package | Runtime | Use it for | +|---|---|---| +| `@questdb/nodejs-client` | Node.js | QWP ingestion and queries over WebSocket, QWP ingestion over UDP, persistent store-and-forward, and the legacy ILP transports | +| `@questdb/browser-client` | Browsers | QWP ingestion and queries over the browser's native WebSocket, with session-cookie authentication. No Node.js dependencies | + +Key capabilities: + +- **Ingestion**: a fluent row API and compiled, type-checked object-row + writers, with automatic table creation, schema evolution, batching, and + acknowledgement tracking. +- **Querying**: SQL with typed bind parameters, results streamed as columnar + batches, DDL and DML execution, cancellation, deadlines, and flow control. +- **One pooled client**: `connectQwpNodeClient()` configures ingestion and + queries from one `ws::` connect string, then hands out pooled senders + (`db.borrowSender()`) and query leases (`db.borrowQuery()`). +- **Failover**: multi-host endpoint lists, automatic reconnect, and replay of + unacknowledged rows. +- **Store-and-forward** (Node.js): a disk journal that keeps accepting rows + while QuestDB is unreachable and survives process restarts. + +:::caution Beta + +QWP support is new in `@questdb/nodejs-client` 5.0.0 and in the first +`@questdb/browser-client` release, and is in beta. Expect the QWP API to change +before it is declared stable, and track the +[client releases](https://github.com/questdb/nodejs-questdb-client/releases). +The ILP transports (`http::`, `https::`, `tcp::`, `tcps::`) are unaffected. -QuestDB offers Node.js developers a dedicated client designed for efficient and -high-performance data ingestion. +::: -:::note No QWP support yet +:::tip Legacy transports -Unlike the other clients in this section, the Node.js client does not speak the -QuestDB Wire Protocol. It ingests over -[ILP](/docs/connect/compatibility/ilp/overview/) using `http::` connect strings, -and queries run over [PGWire](/docs/connect/compatibility/pgwire/nodejs/). The -`http::` examples below are current, not stale. QWP support is planned. +The Node.js `Sender` class still speaks ILP over HTTP and TCP. This page +documents the recommended QWP path. For ILP, see +[ILP transports (legacy)](#ilp-transports-legacy) near the end of this page. ::: -The Node.js client has solid benefits: - -- **Automatic table creation**: No need to define your schema upfront. -- **Concurrent schema changes**: Seamlessly handle multiple data streams with - on-the-fly schema modifications -- **Optimized batching**: Use strong defaults or curate the size of your batches -- **Health checks and feedback**: Ensure your system's integrity with built-in - health monitoring -- **Automatic write retries**: Reuse connections and retry after interruptions +## Requirements -This quick start guide introduces the basic functionalities of the Node.js -client, including setting up a connection, inserting data, and flushing data to -QuestDB. +- **Node.js 20.18.1 or newer** for `@questdb/nodejs-client`. +- **A browser with `WebSocket`, `fetch`, `URL`, `TextEncoder`, and + `TextDecoder`** for `@questdb/browser-client`. +- **QuestDB 10.0.0 or newer**, which serves QWP on the HTTP port (`9000` by + default) at `/write/v4` for ingestion and `/read/v1` for queries. + [Browser connections](#browser-applications) need a QuestDB release newer + than 10.0.1: earlier servers reject upgrades from browsers. If QuestDB is not + running yet, see the [quick start](/docs/getting-started/quick-start/). - +## Installation -:::info +Install the Node.js package: -This page focuses on our high-performance ingestion client, which is optimized for **writing** data to QuestDB. -For retrieving data, we recommend using a [PostgreSQL-compatible Node.js library](/docs/connect/compatibility/pgwire/nodejs/) or our -[HTTP query endpoint](/docs/query/overview/#rest-http-api). +```shell +npm install @questdb/nodejs-client +``` -::: +For browser applications, install the browser package instead: +```shell +npm install @questdb/browser-client +``` -## Requirements +Both packages work with `yarn add` and `pnpm add`. Each exports its complete +API from the package root, ships ES module and CommonJS builds, and bundles +TypeScript declarations. There are no other supported import paths. -- Node.js v16 or newer. -- Assumes QuestDB is running. If it's not, refer to - [the general quick start](/docs/getting-started/quick-start/). +The examples on this page are TypeScript ES modules with top-level `await`. +They also run as plain JavaScript once type annotations are removed. -## Client installation +## Quick start -Install the QuestDB Node.js client via npm: +Connect with one connect string, write two rows, and read them back: -```shell -npm i -s @questdb/nodejs-client +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + // Ingest: borrow a sender, add rows, and close() it to flush the rows and + // return the sender to the pool. The underlying connection stays open. + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.00044) + .at(Date.now(), "ms"); + await sender + .table("trades") + .symbol("symbol", "BTC-USD") + .symbol("side", "sell") + .doubleColumn("price", 39269.98) + .doubleColumn("amount", 0.001) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } + + // Query: borrow a query lease and iterate the result batches. + // QuestDB applies ingested rows asynchronously, so on a first run this + // query can fail with "table does not exist" or return no rows yet. + // See "Read-after-write" below for the polling pattern. + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol, price, amount FROM trades " + + "WHERE symbol = 'ETH-USD' LIMIT 10", + ); + for await (const batch of query) { + for (const [timestamp, symbol, price, amount] of batch.rows()) { + console.log(timestamp, symbol, price, amount); + } + } + await query.completion; + } catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + // QuestDB rejected the SQL: status is the QWP status code. + console.error(`query failed: status=${error.status} ${error.message}`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} ``` -## Authentication +What happens: + +1. `connectQwpNodeClient()` validates every key of the connect string, then + opens one ingestion and one query connection. It rejects if QuestDB is + unreachable. +2. `db.borrowSender()` leases a sender. Rows are staged locally until an + auto-flush threshold is reached or the sender is flushed. `close()` on a + borrowed sender flushes its rows and returns it to the pool. +3. `db.borrowQuery()` leases a query connection. `lease.query()` returns a + query handle that is an async iterable of result batches. `batch.rows()` + yields one array per row. `query.completion` resolves when the server + finishes the query. +4. `db.close()` closes every pooled connection. Pooled senders publish any + remaining rows and wait up to five seconds for QuestDB to acknowledge them. + +The table was created automatically by the first row, so its designated +timestamp column is named `timestamp`. Timestamps come back as `bigint` +microseconds since the Unix epoch; see +[Reading result values](#reading-result-values) for every type. + +### Read-after-write + +**A flush is not a commit, and a commit is not visibility.** QuestDB +acknowledges ingested rows once they are committed to its write-ahead log, and +applies them to the table asynchronously. A query that runs right after +ingestion can therefore fail with `table does not exist` on a first run, or +succeed and return no rows. When your code must read its own writes, poll until +the rows appear, bounded by a deadline: -Passing in a configuration string with basic auth: - -```javascript -const { Sender } = require("@questdb/nodejs-client"); +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, + type QwpClient, +} from "@questdb/nodejs-client"; + +async function countRows(db: QwpClient, sql: string): Promise { + const lease = await db.borrowQuery(); + try { + const query = await lease.query(sql); + let rows = 0; + for await (const batch of query) rows += batch.rowCount; + await query.completion; + return rows; + } finally { + await lease.close(); + } +} -const conf = "http::addr=localhost:9000;username=admin;password=quest;" -const sender = Sender.fromConfig(conf); - ... +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sql = "SELECT * FROM trades WHERE symbol = 'ETH-USD' LIMIT 10"; + const deadline = Date.now() + 10_000; + let rows = 0; + while (rows === 0) { + try { + rows = await countRows(db, sql); + } catch (error) { + // The table may not exist yet: keep polling until the deadline. + if (!(error instanceof QwpEgressQueryError) || Date.now() >= deadline) { + throw error; + } + } + if (rows === 0) { + if (Date.now() >= deadline) throw new Error("rows not visible in time"); + await new Promise((resolve) => setTimeout(resolve, 100)); + } + } + console.log(`visible rows: ${rows}`); +} finally { + await db.close(); +} ``` -Passing via the `QDB_CLIENT_CONF` env var: +Do not replace the poll with a fixed sleep: the apply latency varies with load. -```bash -export QDB_CLIENT_CONF="http::addr=localhost:9000;username=admin;password=quest;" -``` +## Connecting -```javascript -const { Sender } = require("@questdb/nodejs-client"); +There are three ways to create a client in Node.js: +| Entry point | Returns | Use it for | +|---|---|---| +| `connectQwpNodeClient(conf, options?)` | `Promise` | The recommended pooled client for ingestion and queries. Opens the pool minimums and rejects if QuestDB is unreachable. | +| `createQwpNodeClient(conf, options?)` | `QwpClient` | The same pooled client without contacting the server. It connects on `db.connect()` or on the first borrow. | +| `Sender.fromConfig(conf, options?)` | `Promise` | A standalone sender for ingestion only, or for migrating existing ILP code. | +| `connectQwpNodeSender(connection, senderOptions?, sessionOptions?)` | `Promise` | A standalone sender with every column method, configured with typed options instead of a connect string. | -const sender = Sender.fromEnv(); - ... +### Pooled client + +`connectQwpNodeClient()` takes one `ws::` or `wss::` connect string for both +directions. Every `addr` entry is used for ingestion (`/write/v4`) and for +queries (`/read/v1`), and the credentials and TLS keys apply to both: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;sender_pool_max=2;query_pool_max=8;", +); +try { + console.log(db.metrics.senders, db.metrics.queries); +} finally { + await db.close(); +} ``` -When using QuestDB Enterprise, authentication can also be done via REST token. -Please check the [RBAC docs](/docs/security/rbac/#authentication) for more -info. +The `QwpClient` handle has five members: -## Basic insert +| Member | Returns | Purpose | +|---|---|---| +| `borrowSender()` | `Promise` | Lease an exclusive sender. Its `close()` flushes and returns it to the pool. | +| `borrowQuery()` | `Promise` | Lease an exclusive query connection. Its `close()` returns it to the pool. | +| `connect()` | `Promise` | Open the pool minimums. Called for you by `connectQwpNodeClient()`. Safe to retry after a failure. | +| `metrics` | `QwpClientMetrics` | Pool counters (`total`, `available`, `leased`, `creating`, `waiting`) for senders and queries. | +| `close()` | `Promise` | Reject new borrows, close idle connections, cancel active queries, and close the pools. Idempotent. | -Example: inserting executed trades for cryptocurrencies. +Share one `QwpClient` across your application and close it at shutdown. See +[The connection pool](#the-connection-pool) for pool sizing and lease rules. -Without authentication and using the current timestamp. +### Standalone Sender -```javascript -const { Sender } = require("@questdb/nodejs-client") +The `Sender` class predates QWP. Changing its connect string from `http::` to +`ws::` switches it from ILP to QWP while keeping the same row API: -async function run() { - // create a sender using HTTP protocol - const sender = Sender.fromConfig("http::addr=localhost:9000") +```typescript +import { Sender } from "@questdb/nodejs-client"; - // add rows to the buffer of the sender +const sender = await Sender.fromConfig("ws::addr=localhost:9000;"); +try { + // Opens the WebSocket now, so connection errors surface here. + await sender.connect(); await sender .table("trades") .symbol("symbol", "ETH-USD") .symbol("side", "sell") .floatColumn("price", 2615.54) .floatColumn("amount", 0.00044) - .atNow() - - // flush the buffer of the sender, sending the data to QuestDB - // the buffer is cleared after the data is sent, and the sender is ready to accept new data - await sender.flush() - - // close the connection after all rows ingested - // unflushed data will be lost - await sender.close() + .at(Date.now(), "ms"); + await sender.flush(); +} finally { + // Publishes completed rows and waits up to 5 seconds for their ACK. + await sender.close(); } - -run().then(console.log).catch(console.error) ``` -In this case, the designated timestamp will be the one at execution time. Let's -see now an example with an explicit timestamp, custom auto-flushing, and basic -auth. - -```javascript -const { Sender } = require("@questdb/nodejs-client") +`Sender` is ingestion-only. It accepts the complete QWP connect-string +vocabulary, and logs a warning for keys that only the pooled client can apply, +such as `query_pool_max` or `compression`. Its fluent API covers the column +methods that also exist for ILP: `symbol`, `stringColumn`, `booleanColumn`, +`floatColumn` (DOUBLE), `intColumn` (LONG), `timestampColumn`, `arrayColumn`, +`decimalColumn`, and `decimalColumnText`. For the other QuestDB types (UUID, +IPv4, DATE, INT, and more), use a [compiled writer](#compiled-object-row-writers) +through `sender.writer()`, a pooled sender, or `connectQwpNodeSender()`, which +all expose every [column method](#column-methods). -async function run() { - // create a sender using HTTP protocol - const sender = Sender.fromConfig( - "http::addr=localhost:9000;username=admin;password=quest;auto_flush_rows=100;auto_flush_interval=1000;", - ) +`connectQwpNodeSender()` builds a standalone `QwpSender` from typed options. +Its first argument takes the full ingestion URL, and credentials as an +`authorization` header value such as `` `Bearer ${token}` ``: - // Calculate the current timestamp. You could also parse a date from your source data. - const timestamp = Date.now() +```typescript +import { connectQwpNodeSender } from "@questdb/nodejs-client"; - // add rows to the buffer of the sender +const sender = await connectQwpNodeSender( + { url: "ws://localhost:9000/write/v4" }, + { autoFlushRows: 5_000, autoFlushIntervalMs: 1_000 }, +); +try { await sender .table("trades") .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.00044) - .at(timestamp, "ms") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .uuidColumn("order_id", "9f1c96b2-54b8-4d85-bb24-e82c6f1ac120") + .at(Date.now(), "ms"); + await sender.flush(); +} finally { + await sender.close(); +} +``` - // add rows to the buffer of the sender - await sender - .table("trades") - .symbol("symbol", "BTC-USD") - .symbol("side", "sell") - .floatColumn("price", 39269.98) - .floatColumn("amount", 0.001) - .at(timestamp, "ms") +### Environment variable + +Keep credentials out of source code by putting the connect string in the +`QDB_CLIENT_CONF` environment variable: + +```bash +export QDB_CLIENT_CONF="wss::addr=db.example.com:9000;token=YOUR_TOKEN;" +``` - // flush the buffer of the sender, sending the data to QuestDB - // the buffer is cleared after the data is sent, and the sender is ready to accept new data - await sender.flush() +`Sender.fromEnv()` reads the variable. The pooled client takes the string +directly: - // close the connection after all rows ingested - // unflushed data will be lost - await sender.close() +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const conf = process.env.QDB_CLIENT_CONF; +if (!conf) throw new Error("QDB_CLIENT_CONF is not set"); +const db = await connectQwpNodeClient(conf); +try { + // borrow senders and query leases +} finally { + await db.close(); } +``` + +### Connect string syntax + +A QWP connect string has the form `schema::key=value;key=value;`: + +- **Schema**: `ws` (plain WebSocket) or `wss` (WebSocket over TLS). Both + default to port `9000` when `addr` omits the port. +- **`addr`**: `host[:port]`. List several endpoints for failover, either + comma-separated (`addr=a:9000,b:9000`) or by repeating the key. Enclose IPv6 + addresses in brackets: `addr=[::1]:9000`. +- **Keys** are lowercase and case-sensitive. An unrecognized key fails with + `unknown configuration key: `. Legacy ILP keys such as `retry_timeout` + or `init_buf_size` fail with a hint that names the QWP replacement. +- **Values** end at `;`. Double a semicolon to include it in a value: + `password=p;;ssw;;rd` sets the password to `p;ssw;rd`. The trailing `;` is + optional. + +The JavaScript parser differs from some other clients in two places: + +- `auto_flush_rows` and `auto_flush_interval` take `0`, not `off`, to disable + a trigger. `auto_flush=off` disables auto-flushing entirely. +- Size values accept the single-letter suffixes `k`, `m`, `g`, and `t` + (`sf_max_total_bytes=10g`). The two-letter forms `kb`, `mb`, and `gb` are + rejected. + +For every key and its default, see the +[connect string reference](/docs/connect/clients/connect-string/) and the +[configuration reference](#configuration-reference) at the end of this page. + +### Programmatic options + +Callbacks, custom agents, and other settings a string cannot express go in the +second argument. The connect string is validated in full first; when both set +the same option, the typed value wins: -run().then(console.log).catch(console.error) +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + // QwpSenderOptions for every pooled sender + sender: { awaitServerAck: true }, + // Ingestion callbacks and replay settings + ingressSession: { + onSenderError: (error) => + console.error("rejected batch", error.category, error.serverMessage), + }, + // Query session defaults + egressSession: { queryTimeoutMs: 30_000 }, + // Egress-only routing and compression + egress: { compression: "zstd" }, + // Pool sizes and timeouts + pool: { senderPoolMax: 2, queryPoolMax: 8 }, +}); +await db.close(); ``` -As you can see, both events now are using the same timestamp. We recommended to -use the original event timestamps when ingesting data into QuestDB. Using the -current timestamp hinder the ability to deduplicate rows which is -[important for exactly-once processing](/docs/connect/compatibility/ilp/overview/#exactly-once-delivery-vs-at-least-once-delivery). +The other sections are `webSocket` (connection settings shared by both +directions, such as `agent` or `connectTimeoutMs`) and `storeAndForward` +(journal settings, see [Store-and-forward](#store-and-forward)). + +`Sender.fromConfig()` takes `{ log, agent, qwp }` as its second argument, +where `qwp` has the sections `webSocket`, `session` (the equivalent of +`ingressSession`), `sender`, and `udp`. -## Decimal insertion +:::caution A typed `reconnect` object replaces the connect-string keys -:::note -Decimal columns are available with ILP protocol version 3 (QuestDB v9.2.0+ and NodeJS client v4.2.0+). +`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces +the whole reconnect policy parsed from `reconnect_*` keys, and +`egressSession.reconnect` replaces the policy parsed from `failover*` keys. +Fields you leave out of the object take the built-in defaults, not the values +from the connect string. When you supply the object, for example to register +`onEvent`, set every bound you rely on in it, such as `maxDurationMs`. -HTTP/HTTPS connections negotiate this automatically (`protocol_version=auto`), while TCP/TCPS connections must opt in explicitly (for example `tcp::...;protocol_version=3`). Once on v3, you can choose between the textual helper and the binary helper. ::: -:::caution -QuestDB does not auto-create decimal columns. Define them ahead of ingestion with -`DECIMAL(precision, scale)` so the server knows how many digits to store, as explained in the -[decimal data type](/docs/query/datatypes/decimal/#creating-tables-with-decimals) guide. +## Authentication and TLS + +QWP authenticates on the WebSocket upgrade request, before any data is +exchanged. The credential and TLS keys apply to both ingestion and queries. + +### Token (Enterprise, recommended) + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); +const db = await connectQwpNodeClient( + `wss::addr=db.example.com:9000;token=${token};`, +); +await db.close(); +``` + +The token is sent as an `Authorization: Bearer` header on every ingestion and +query upgrade. A [REST token](/docs/security/rbac/#authentication) and an OIDC +access token both use `token`. It cannot be combined with +`username`/`password`. + +### HTTP basic auth + +```text +wss::addr=db.example.com:9000;username=admin;password=quest; +``` + +`user` and `pass` are accepted aliases. Both halves must be present, and the +username cannot contain `:`. + +### TLS + +The `wss` schema enables TLS and verifies the server certificate against the +system trust store. Two keys adjust verification, and both are rejected on a +plain `ws` string: + +- `tls_roots=/path/to/ca.pem` trusts the CA certificates in a PEM file instead + of the system store. The Node.js client accepts PEM only: + `tls_roots_password` and PKCS#12 or JKS stores are rejected. Export the CA + certificates to PEM first. +- `tls_verify=unsafe_off` disables certificate verification. Use it only in + development. It cannot be combined with `tls_roots`. + +To route the connection through an HTTP or SOCKS proxy, pass an agent such as +`https-proxy-agent` in `webSocket.agent` (or `qwp.webSocket.agent` on a +`Sender`). A custom agent owns certificate verification, so it cannot be +combined with `tls_verify` or `tls_roots`. + +Two deadlines bound connection setup: `connect_timeout` covers DNS and the +TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and +authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits +`connect_timeout` when only the latter is set. A timeout fails with +`QwpUpgradeError`, whose `timeoutPhase` is `connect` or `authentication`. + +### Unsupported authentication paths + +| Path | Status | Workaround | +|---|---|---| +| OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider and has no callback to refresh a token. | Obtain an access token from your identity provider, pass it as `token=...`, and create a new client before the token expires. See [OpenID Connect](/docs/security/oidc/). | +| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. | Close the client and create a new one with the new token. | +| Mutual TLS (client certificates) | Not supported. QuestDB does not negotiate client certificates. | Use token or basic authentication over `wss`. | +| ILP JWK authentication | Not available for QWP. `auth`, `jwk`, `token_x`, and `token_y` are rejected on `ws`/`wss`. | Use token or basic authentication. | + +### Production example: TLS, token, and multiple hosts + +A typical Enterprise deployment combines `wss`, a token, and several hosts: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); + +const db = await connectQwpNodeClient( + "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + + `token=${token};` + + "tls_roots=/etc/ssl/questdb-ca.pem;", + { + // Serve queries from replicas. Keep this out of the connect string: there, + // `target` also applies to ingestion, which must reach the primary. + egress: { target: "replica" }, + }, +); +try { + // borrow senders and query leases +} finally { + await db.close(); +} +``` + +## The connection pool + +The pooled client keeps two elastic pools: one of senders and one of query +connections. Each pool opens its minimum on `connect()`, grows on demand up to +its maximum, and a housekeeper closes connections that stay idle too long or +exceed their maximum lifetime, never going below the minimum. + +### Borrowing a sender + +A borrowed sender belongs to the borrower until its `close()` returns it: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + for (const [symbol, price] of [ + ["ETH-USD", 2615.54], + ["BTC-USD", 39269.98], + ] as const) { + await sender + .table("trades") + .symbol("symbol", symbol) + .symbol("side", "buy") + .doubleColumn("price", price) + .doubleColumn("amount", 0.1) + .at(Date.now(), "ms"); + } + } finally { + // Flushes the rows and returns the sender. Does not wait for the ACK. + await sender.close(); + } +} finally { + await db.close(); +} +``` + +A long-running producer can keep its borrow for its whole lifetime and call +`flush()` between batches. Size `sender_pool_max` to the number of producers +that hold a sender at the same time. + +:::note Pooled sender close semantics + +`close()` on a borrowed sender flushes completed rows, discards an unfinished +row with a warning, and returns the sender to the pool. It does not close the +WebSocket and does not wait for QuestDB to acknowledge the rows. To confirm +delivery before returning the sender, use +[`flushAndGetSequence()` and `waitForAcknowledged()`](#awaiting-acknowledgements). + +When a borrowed sender's `close()` fails, the pool discards that sender and +opens a new one for the next borrow. Because QuestDB reports rejected batches +asynchronously, a sender can fail after its `close()` already succeeded: the +error then surfaces on the `flush()` or `close()` of the next borrower, and the +pool replaces the sender after that. See [Ingestion errors](#ingestion-errors). + ::: -### Text literal (easy to use) +### Borrowing a query lease + +A query lease runs one query at a time. For concurrent queries, borrow one lease +per query, up to `query_pool_max`: ```typescript -import { Sender } from "@questdb/nodejs-client"; +import { connectQwpNodeClient, type QwpQueryLease } from "@questdb/nodejs-client"; -async function runDecimalsText() { - const sender = await Sender.fromConfig( - "tcp::addr=localhost:9009;protocol_version=3", +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); + +async function countBySymbol(lease: QwpQueryLease, symbol: string) { + const query = await lease.query( + "SELECT count() FROM trades WHERE symbol = $1", + { binds: (binds) => binds.setVarchar(0, symbol) }, ); + let count = 0n; + for await (const batch of query) count = batch.get(0, 0) as bigint; + await query.completion; + return count; +} - await sender - .table("fx") - .symbol("pair", "EURUSD") - .decimalColumnText("mid", "1.234500") // keeps trailing zeros - .atNow(); +try { + const [a, b] = await Promise.all([db.borrowQuery(), db.borrowQuery()]); + try { + // Two leases, two WebSockets: the queries run concurrently. + const [eth, btc] = await Promise.all([ + countBySymbol(a, "ETH-USD"), + countBySymbol(b, "BTC-USD"), + ]); + console.log({ eth, btc }); + } finally { + await Promise.all([a.close(), b.close()]); + } +} finally { + await db.close(); +} +``` + +Starting a second query on a lease while one is still active throws +`a QWP query is already active on this connection`. Always close a lease in +`finally`: an unreturned lease holds its connection until `db.close()`. + +### Pool settings + +| Key | Default | Purpose | +|---|---|---| +| `sender_pool_min` | `1` | Senders kept open even when idle. `0` lets the pool close them all. | +| `sender_pool_max` | `4` | Maximum senders the pool opens. | +| `query_pool_min` | `1` | Query connections kept open even when idle. | +| `query_pool_max` | `4` | Maximum query connections, which also caps concurrent queries. | +| `acquire_timeout_ms` | `5000` | How long a borrow waits when the pool is at its maximum, before rejecting with `QwpPoolAcquireTimeoutError`. | +| `idle_timeout_ms` | `60000` | Idle time before an excess connection is closed. `0` keeps idle connections. | +| `max_lifetime_ms` | `1800000` | Age at which an idle connection is recycled. `0` disables recycling. | +| `housekeeper_interval_ms` | `5000` | How often the housekeeper checks for idle and over-age connections. Minimum `100`. | +| `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | +| `lazy_connect` | `off` | Start without connecting. See below. | + +The typed equivalents live in the `pool` section of the second argument +(`senderPoolMin`, `acquireTimeoutMs`, `housekeepingIntervalMs`, and so on). +When creating a new pooled connection fails, the borrow rejects with +`QwpPoolResourceError`, whose `cause` holds the connection error. + +### Starting while QuestDB is down + +`connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set +`lazy_connect=on` to start regardless: senders connect in the background and +buffer rows in memory until QuestDB is reachable, and the query pool stays +empty until the first query. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// Resolves immediately, even if QuestDB is not running yet. +const db = await connectQwpNodeClient("ws::addr=localhost:9000;lazy_connect=on;"); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +`lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, +and rejects an explicit conflicting value. A query borrowed while QuestDB is +still down rejects with `QwpPoolResourceError`. + +## Data ingestion + +### General usage pattern + +A sender is not safe for concurrent producers: the row in progress is shared +state, so borrow one sender per producer (see [Concurrency](#concurrency)). +1. Borrow a sender with `db.borrowSender()`, or create a + [standalone `Sender`](#standalone-sender). +2. Call `table(name)` to start a row. +3. Add values with the [column methods](#column-methods), such as + `symbol(name, value)` and `doubleColumn(name, value)`. To store a NULL, pass + `null` or `undefined`, or skip the column (see [Null values](#null-values)). +4. Close the row with `at(timestamp, unit)` or `atNow()`, and `await` the + returned promise. It rejects if an auto-flush triggered by the row fails. +5. Repeat from step 2, and call `flush()` to send staged rows. +6. `close()` the sender when done. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.25) + .at(Date.now(), "ms"); + } catch (error) { + // An invalid value throws a TypeError or RangeError. The row in progress + // is discarded; rows completed earlier stay staged. + console.error("row rejected:", error); + } + await sender.flush(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +Tables and columns are created automatically, with the column types listed +below. Table and column names are validated locally with QuestDB's rules +(at most 127 UTF-8 bytes by default, see `max_name_len`), and column names are +case-insensitive: the first spelling used is kept. + +When a column method or `at()` rejects a value, the sender discards the whole +row in progress, including its table, so a half-built row never reaches +QuestDB. The next row must start with `table()` again; a column method called +before that throws `table name must be set before adding columns`. +`cancelRow()` discards a row in progress without an error, and `reset()` also +drops every row staged since the last flush. + +### Column methods + +These methods are available on pooled senders and on senders from +`connectQwpNodeSender()` and `connectQwpBrowserSender()`. Each creates the +listed column type when the column does not exist yet: + +| Method | QuestDB type created | Accepted values | +|---|---|---| +| `symbol(name, value)` | SYMBOL | Any value, converted with `String()` | +| `stringColumn(name, value)` | VARCHAR | `string` | +| `booleanColumn(name, value)` | BOOLEAN | `boolean` | +| `byteColumn(name, value)` | BYTE | Integer `number` from -128 to 127 | +| `shortColumn(name, value)` | SHORT | Integer `number` from -32768 to 32767 | +| `int32Column(name, value)` | INT | 32-bit integer `number`. `-2147483648` stores NULL | +| `longColumn(name, value)`, `intColumn(name, value)` | LONG | Safe-integer `number` or `bigint`. `-9223372036854775808n` stores NULL | +| `float32Column(name, value)` | FLOAT | `number` | +| `doubleColumn(name, value)`, `floatColumn(name, value)` | DOUBLE | `number` | +| `timestampColumn(name, value, unit)` | TIMESTAMP, or TIMESTAMP_NS with unit `"ns"` | Integer `number` or `bigint`. Unit `"us"` (default), `"ms"`, or `"ns"`; `"ns"` requires a `bigint` | +| `dateColumn(name, value)` | DATE | Epoch milliseconds as `number` or `bigint` | +| `charColumn(name, value)` | CHAR | One-character `string` (a single UTF-16 code unit) | +| `binaryColumn(name, value)` | BINARY | `Uint8Array`, copied when staged | +| `uuidColumn(name, value)` | UUID | Canonical UUID `string`, or 16 bytes in canonical big-endian order | +| `long256Column(name, w0, w1, w2, w3)` | LONG256 | Four 64-bit `bigint` words, least significant first | +| `ipv4Column(name, value)` | IPv4 | Dotted-quad `string` or packed 32-bit `number`. `0.0.0.0` stores NULL | +| `geohashColumn(name, bits, precisionBits)` | GEOHASH | Raw bits as `bigint`, precision from 1 to 60 bits | +| `decimalColumnText(name, value)` | DECIMAL(76, scale) | Decimal `string` or `number`. The scale comes from the literal | +| `decimalColumn(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, or big-endian two's-complement `Int8Array` | +| `decimal64Column(name, unscaled, scale)` | DECIMAL(18, scale) | Unscaled `bigint`, scale up to 18 | +| `decimal128Column(name, unscaled, scale)` | DECIMAL(38, scale) | Unscaled `bigint`, scale up to 38 | +| `decimal256Column(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, scale up to 76 | +| `arrayColumn(name, value)` | DOUBLE[], DOUBLE[][], ... | Nested `number` arrays of uniform shape, 1 to 32 dimensions | +| `longArrayColumn(name, value)` | LONG[] | Encoded for protocol parity, but current QuestDB servers reject LONG array ingestion | + +Names that differ from what you might expect: + +- `floatColumn()` and `intColumn()` write 64-bit DOUBLE and LONG. Use + `float32Column()` and `int32Column()` for FLOAT and INT. +- There is no `nullColumn()` or `setNull()`. Pass `null` or `undefined`, or + skip the column. +- Arrays use `arrayColumn()`. `doubleArray()` is a + [compiled writer](#compiled-object-row-writers) field, not a sender method. +- `geohashColumn()` takes raw bits only. Base-32 geohash text is accepted by a + compiled writer's `geohash()` field. + +The standalone `Sender` class exposes only `symbol`, `stringColumn`, +`booleanColumn`, `floatColumn`, `intColumn`, `timestampColumn`, `arrayColumn`, +`decimalColumn`, and `decimalColumnText`. Its `writer()` method supports every +type. + +A column's type is fixed by the first value a sender stages for it. Writing a +different type to the same column later throws +`column type mismatch for ''`. If the table already exists with a +different column type, QuestDB rejects the batch asynchronously; see +[Ingestion errors](#ingestion-errors). + +### Null values + +To store NULL, pass `null` or `undefined` to any column method, or leave the +column out of the row. All three have the same effect: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + const trade: { side?: string; amount?: number } = { amount: 0.011 }; + await sender + .table("trades") + .symbol("symbol", "BTC-USD") + .symbol("side", trade.side) // undefined: stored as NULL + .doubleColumn("price", 39269.98) + .doubleColumn("amount", trade.amount) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +- An omitted column is not created on a table that lacks it: a NULL carries no + type to infer from. +- The column name is still validated when the value is nullish. +- Rows that already exist in a batch, or rows added later, get NULL for any + column they do not set. +- INT, LONG, and DATE reserve their minimum values as NULL: writing + `-2147483648` to INT or `-9223372036854775808n` to LONG or DATE stores NULL. + IPv4 treats `0.0.0.0` as NULL. +- A row where every value is nullish is still sent over WebSocket and stored + with NULL in every column. To drop such a row instead, call `cancelRow()` + before closing it. Over UDP, `atNow()` rejects such a row while the sender + knows no non-null column for the table. + +### Designated timestamp + +The [designated timestamp](/docs/concepts/designated-timestamp/) controls +partitioning and ordering. Set it when closing the row: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + // Milliseconds, for example from Date.now() + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + + // Microseconds are the default unit + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .doubleColumn("price", 2615.55) + .doubleColumn("amount", 0.2) + .at(BigInt(Date.now()) * 1000n); + + // Server-assigned timestamp + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.56) + .doubleColumn("amount", 0.1) + .atNow(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +`at(value, unit)` accepts an integer `number` or a `bigint` with unit `"us"` +(the default), `"ms"`, or `"ns"`. Nanoseconds require a `bigint`, because epoch +nanoseconds exceed the safe integer range. When the table does not exist yet, +`"ns"` creates a `TIMESTAMP_NS` designated timestamp and the other units create +a microsecond `TIMESTAMP`. An auto-created designated timestamp column is named +`timestamp`. + +`atNow()` leaves the timestamp to QuestDB, which assigns it when the row +arrives. Rows replayed after a reconnect are stamped with the replay time. +Prefer event timestamps from your source data: they keep rows in event order and +make [deduplication](/docs/concepts/deduplication/) possible, which is +[required for exactly-once delivery](/docs/concepts/delivery-semantics/). + +Other timestamp columns use `timestampColumn(name, value, unit)` with the same +units. For converting dates and strings, see +[Date to timestamp conversion](/docs/connect/clients/date-to-timestamp-conversion/). + +### Arrays + +`arrayColumn()` takes nested `number` arrays and creates a `DOUBLE` array column +with the same number of dimensions: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("order_book") + .symbol("symbol", "BTC-USD") + // shape [2, N]: row 0 holds prices, row 1 holds sizes + .arrayColumn("bids", [ + [64901.6, 64901.5, 64901.4], + [3.02, 0.06, 1.2], + ]) + .arrayColumn("asks", [ + [64901.7, 64901.8, 64901.9], + [1.54, 0.21, 2.5], + ]) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +Every sub-array at the same depth must have the same length, and arrays may +have 1 to 32 dimensions. Only DOUBLE arrays can be ingested: +`longArrayColumn()` exists for protocol parity, but current servers reject it +with `long arrays are not supported, only double arrays`. Query results return +arrays as `{ dimensions, values }`; see +[Reading result values](#reading-result-values). + +### Decimals + +Create decimal columns ahead of time with the precision you need. QWP can +create them automatically, but it picks the maximum precision of the wire +width (18, 38, or 76 digits). See +[decimal data type](/docs/query/datatypes/decimal/#creating-tables-with-decimals). + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const ddl = await lease.query( + "CREATE TABLE IF NOT EXISTS trade_fees (" + + "timestamp TIMESTAMP, symbol SYMBOL, " + + "settled_price DECIMAL(18, 2), commission DECIMAL(18, 4)" + + ") TIMESTAMP(timestamp) PARTITION BY DAY", + ); + await ddl.completion; + } finally { + await lease.close(); + } + + const sender = await db.borrowSender(); + try { + await sender + .table("trade_fees") + .symbol("symbol", "ETH-USD") + .decimal64Column("settled_price", 261554n, 2) // 2615.54 + .decimalColumnText("commission", "0.0750") // keeps the literal's scale + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +- `decimalColumnText()` takes a decimal string, scientific notation included + (`"1.5e-3"`), and preserves the literal's scale, including trailing zeros. + Passing a `number` works, but JavaScript drops trailing zeros when formatting. +- `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` + or as big-endian two's-complement bytes in an `Int8Array`. +- `decimal64Column()`, `decimal128Column()`, and `decimal256Column()` take an + unscaled `bigint` and select the wire width directly. + +Scale rules: + +- The first value staged for a decimal column fixes its scale until the next + flush. Later values are rescaled exactly (`"2.50"` becomes `2.5` at scale 1), + and a value that would lose digits throws a `RangeError`, such as `"1.25"` + at scale 1. +- When QWP creates the column, the first value's scale becomes the column's + scale. +- QuestDB converts each value to the table column's scale when no digits are + lost: `"2615.5400"` is stored as `2615.54` in a `DECIMAL(18, 2)` column. A + value that would lose digits, such as `0.0015` for `DECIMAL(18, 2)`, fails + the whole batch with a terminal `schema-mismatch` rejection. Stage values + with the column's scale. + +### Compiled object-row writers + +When your data is already a stream of objects with one shape, compile a writer +for the table once. The writer validates each complete row before staging it, +and TypeScript checks every row against the schema: + +```typescript +import { + connectQwpNodeClient, + designatedTimestamp, + double, + QwpWriterRowError, + symbol, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + const trades = sender.writer("trades", { + symbol: symbol(), + side: symbol(), + price: double(), + amount: double(), + timestamp: designatedTimestamp("ms"), + }); + + await trades.row({ + symbol: "ETH-USD", + side: "sell", + price: 2615.54, + amount: 0.00044, + timestamp: Date.now(), + }); + + // Arrays, iterables, and async iterables. Absent fields store NULL. + await trades.rows([ + { symbol: "BTC-USD", side: "buy", price: 39269.98, timestamp: Date.now() }, + ]); + } catch (error) { + if (!(error instanceof QwpWriterRowError)) throw error; + // Names the table, the column, and the zero-based row index for rows(). + console.error(error.message); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +A rejected row is never partly staged, and rows accepted before a failing row +in `rows()` stay staged. Unknown keys and type mismatches raise +`QwpWriterRowError`. Writers apply the sender's normal auto-flush, transaction, +and acknowledgement settings. `writer()` works on pooled senders and on the +standalone `Sender` over `ws`, `wss`, or `udp`; an ILP `Sender` throws. + +Schema fields: + +| Field | QuestDB type | Row value | +|---|---|---| +| `symbol()` | SYMBOL | `string` | +| `varchar()` | VARCHAR | `string` | +| `char()` | CHAR | One-character `string` | +| `bool()` | BOOLEAN | `boolean` | +| `byte()`, `short()` | BYTE, SHORT | `number` | +| `int32()` | INT | `number` | +| `int64()`, `long()` | LONG | `bigint` | +| `float32()` | FLOAT | `number` | +| `float64()`, `double()` | DOUBLE | `number` | +| `timestamp(unit)` | TIMESTAMP or TIMESTAMP_NS | `number` or `bigint`; `"ns"` requires `bigint` | +| `designatedTimestamp(unit)` | designated TIMESTAMP | As `timestamp(unit)`, required in every row. At most one per schema | +| `date()` | DATE | Epoch milliseconds | +| `binary()` | BINARY | `Uint8Array` | +| `uuid()` | UUID | Canonical UUID `string`, 16 big-endian bytes, or `{ low, high }` | +| `long256()` | LONG256 | Unsigned 256-bit `bigint`, `0x` hex text, four little-endian words, or `{ words }` | +| `ipv4()` | IPv4 | Dotted-quad `string` or packed `number` | +| `geohash(precisionBits)` | GEOHASH | Raw bits, base-32 text of `precisionBits / 5` characters, or `{ bits, precisionBits }` | +| `decimal64(scale)`, `decimal128(scale)`, `decimal256(scale)` | DECIMAL | Unscaled `bigint`, decimal text, `number`, or `{ unscaled, scale }` | +| `doubleArray()` | DOUBLE[] | Nested `number` arrays, or `{ dimensions, values }` | +| `longArray()` | LONG[] | Encoded for parity; current servers reject LONG arrays | + +LONG fields take `bigint` so they never lose precision. The object forms +(`{ low, high }`, `{ words }`, `{ bits, precisionBits }`, `{ unscaled, scale }`, +`{ dimensions, values }`) match what [query results](#reading-result-values) +return, so a queried value can be written back unchanged. A writer's decimal +field rescales values to the field's scale and rejects a value that would need +rounding. + +### Flushing + +Rows are staged in memory until a flush publishes them. Auto-flush is on by +default and flushes after the row that crosses the first threshold: + +| Trigger | Default | Connect-string key | Typed option | +|---|---|---|---| +| Row count | 1,000 rows | `auto_flush_rows` | `autoFlushRows` | +| Time since the last flush | 100 ms | `auto_flush_interval` | `autoFlushIntervalMs` | +| Estimated buffered bytes | Disabled | `auto_flush_bytes` | `autoFlushBytes` | + +The interval is checked when a row is added. There is no background timer, so +call `flush()` after a burst of rows, or rows staged before an idle period wait +for the next row. `auto_flush=off` disables all triggers. `auto_flush_bytes` is +clamped to 90% of the batch size the server advertises. + +What `flush()` waits for depends on the ingestion mode: + +| Mode | Enabled by | `flush()` resolves when | During an outage | +|---|---|---|---| +| Memory (default) | Neither of the others | The batch is written to the WebSocket, or queued for replay | `flush()` and auto-flushing `at()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | +| Background memory | `initial_connect_retry=async`, or `lazy_connect=on` on the pooled client | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | +| Store-and-forward (Node.js) | `sf_dir` | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | + +In every mode, `flush()` does not wait for QuestDB to acknowledge the rows, +unless you set `awaitServerAck`. Unacknowledged batches are kept and replayed +after a reconnect. See [Awaiting acknowledgements](#awaiting-acknowledgements) +and [Store-and-forward](#store-and-forward). + +**Backpressure.** The in-memory replay queue is capped at 128 MiB. When it is +full, publishing waits up to 30 seconds for acknowledgements to free space, +then rejects with `QwpMemoryReplayAppendTimeoutError`. Tune the cap with +`sf_max_total_bytes` and the wait with `sf_append_deadline_millis`; without +`sf_dir` they size the memory queue. Watch `sender.metrics.ingress` +(`memoryReplayUsedBytes`, `totalMemoryReplayBackpressureStalls`) to detect +backpressure before it blocks. A single row larger than the server's batch +limit is rejected before it is sent, with `QwpBatchTooLargeError`. + +**Closing.** `close()` on a standalone sender publishes completed rows and waits +up to `close_flush_timeout_millis` (5 seconds by default) for their +acknowledgement. `0` or a negative value skips the wait. An unfinished row is +discarded with a warning. On a borrowed sender, `close()` flushes and returns +the sender to the pool without waiting; see +[Borrowing a sender](#borrowing-a-sender). + +### Awaiting acknowledgements + +QuestDB acknowledges ingested batches asynchronously. Every published frame gets +a sequence number, and the acknowledgement watermark is cumulative, so waiting +for one sequence also covers every earlier one: + +```typescript +import { + connectQwpNodeClient, + QwpIngressAckTimeoutError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + + const sequence = await sender.flushAndGetSequence(); + // Rejects with the server's error if QuestDB rejected the batch. + await sender.waitForAcknowledged(sequence, 10_000); + } catch (error) { + if (error instanceof QwpIngressAckTimeoutError) { + // Not acknowledged in time. The rows are still pending, not lost. + console.warn("ACK timeout at", error.acknowledgedSequence); + } else { + throw error; + } + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +| Member | Returns | +|---|---| +| `flushAndGetSequence()` | Publishes staged rows and resolves with the highest sequence (`bigint`) this call published, or `-1n` when there was nothing to publish. | +| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout, without closing the sender, or with the server's rejection. | +| `acknowledgedSequence` | The highest acknowledged sequence, or `-1n`. | +| `publishedSequence` | The highest published sequence, or `-1n`. | + +To make every `flush()` wait for its acknowledgement, set `awaitServerAck`: +`connectQwpNodeClient(conf, { sender: { awaitServerAck: true } })`, or +`{ qwp: { sender: { awaitServerAck: true } } }` for a standalone `Sender`. A +server rejection then rejects `flush()` itself, with `QwpIngressNackError`. + +Acknowledgement is not required for delivery: unacknowledged batches are +replayed after a reconnect, and a standalone sender waits for them on +`close()`. Wait for acknowledgements when your application must know that +QuestDB accepted the rows, for example before committing a source offset. If +the process exits before the acknowledgement, rows still in memory are lost; +use [store-and-forward](#store-and-forward) to keep them across restarts. + +### Transactions + +By default each flushed batch is committed on its own. With transactions on, +auto-flushed batches stay in an open server-side transaction until you commit: + +```typescript +import { Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig( + "ws::addr=localhost:9000;transaction=on;auto_flush_rows=10000;", +); +try { + await sender.connect(); + for (let i = 0; i < 50_000; i++) { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", i % 2 === 0 ? "buy" : "sell") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.01) + .at(Date.now(), "ms"); + } + // Commits every auto-flushed batch plus the staged rows. await sender.flush(); +} finally { await sender.close(); } ``` -`decimalColumnText` accepts strings or numbers. String literals go through `validateDecimalText` and are written verbatim with the `d` suffix, so every digit (including trailing zeros or exponent form) is preserved. Passing a number is convenient, but JavaScript’s default formatting will drop insignificant zeros. +- The transaction is atomic per table. A flush that spans several tables + commits each table separately. +- `flush()` commits. Pooled senders also have `commit()`, an alias of + `flush()`. The typed option is `transactional: true`. +- Closing without committing rolls the open transaction back, with a warning. +- QuestDB does not acknowledge the deferred batches until the commit, so + `waitForAcknowledged()` for a sequence inside an open transaction waits for + the commit. + +### Store-and-forward + +In the default memory mode, unacknowledged rows are lost if the process +exits. Setting `sf_dir` (Node.js only) turns on a disk journal instead: every +batch is appended to the journal before it is sent, a background drainer sends +it in order, and acknowledged segments are deleted. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;" + + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + + "sf_durability=append;initial_connect_retry=async;", +); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + // Resolves once the rows are in the journal, even if QuestDB is down. + await sender.flush(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +With a journal, the sender keeps accepting rows while QuestDB is unreachable, +until `sf_max_total_bytes` (10 GiB) is full. It retries the connection +indefinitely once it has connected, and a new sender opened on the same +directory replays what the previous process left behind. + +- **Layout.** A standalone `Sender` journals into `/`. A + pooled client uses one directory per pooled sender: + `/-0`, `/-1`, and so on. `sender_id` + defaults to `default` and may contain letters, digits, `_`, and `-`. Give + every process its own `sender_id`; a second live process on the same + directory fails with `QwpReplayStoreLockedError`. +- **Durability.** `sf_durability=memory` (the connect-string default) relies on + the operating system to write the journal, which survives a process crash but + not a power loss. `periodic` checkpoints in the background every + `sf_sync_interval_millis` (5 seconds). `append` makes every append durable + before `flush()` resolves. +- **Backpressure.** When the journal is full, publishing waits up to + `sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, + then rejects with `QwpReplayStoreAppendTimeoutError`. +- **Startup.** `initial_connect_retry=async` lets a sender start while QuestDB + is down. With the default `off`, the first connection must succeed. +- **Orphans.** With `drain_orphans=on`, a sender also adopts and drains + journals left under the same `sf_dir` by processes that crashed, up to + `max_background_drainers` (4) at a time. + +A frame appended to the journal but not acknowledged before a crash is sent +again, so delivery is at least once: + + + +:::warning Do not share a journal directory with a Java client + +The Node.js client locks journal directories with its own lock files, which the +Java client does not see. Never point a running Java client and a running +Node.js client at the same directory. The journal format is shared, so a +directory written by one can be opened by the other after the first has +closed it. + +::: + +For all tuning options, see +[Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/) +and the [store-and-forward keys](/docs/connect/clients/connect-string/#sf-keys). + +### Durable acknowledgement + +:::note Enterprise + +Durable acknowledgement requires QuestDB Enterprise with primary replication +configured. + +::: + +By default QuestDB acknowledges a batch when it is committed to the primary's +write-ahead log. With `request_durable_ack=on`, the acknowledgement watermark +advances only after the batch is uploaded to the replication object store, so +`waitForAcknowledged()` confirms durable upload: + +```text +wss::addr=db.example.com:9000;token=YOUR_TOKEN;request_durable_ack=on; +``` + +To make every flush wait for durability, add the typed option +`{ sender: { awaitDurableAck: true } }`. If the server does not support durable +acknowledgement, connecting fails with `QwpDurableAckUnavailableError`. -### Binary form (high throughput) +### Fire-and-forget UDP + +The Node.js `Sender` can send rows as UDP datagrams, for metrics where +occasional loss is acceptable: ```typescript +import { Sender } from "@questdb/nodejs-client"; + const sender = await Sender.fromConfig( - "tcp::addr=localhost:9009;protocol_version=3", + "udp::addr=localhost:9007;max_datagram_size=1400;", ); +try { + await sender.connect(); + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); +} finally { + await sender.close(); +} +``` + +UDP has no authentication, TLS, acknowledgements, transactions, reconnect, or +store-and-forward, and it is not available in browsers. The server's UDP +receiver is disabled by default; enable it with +[`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver). The default port is +`9007`. `max_datagram_size` (1400 bytes by default) must fit your network path; +a row that cannot fit a datagram fails with `QwpUdpDatagramTooLargeError`. +`multicast_ttl` sets the multicast time-to-live. + +## Querying + +Queries run on a lease borrowed from the pooled client. One lease runs one +query at a time, so borrow one lease per concurrent query. + +### Running a SELECT + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol, price, amount FROM trades " + + "WHERE symbol = $1 AND price > $2 LIMIT 100", + { + binds: (binds) => binds.setVarchar(0, "ETH-USD").setDouble(1, 2000), + timeoutMs: 30_000, + }, + ); + for await (const batch of query) { + for (const [timestamp, symbol, price, amount] of batch.rows()) { + console.log(timestamp, symbol, price, amount); + } + } + const completion = await query.completion; + console.log("rows:", completion.kind === "result-end" && completion.totalRows); + } catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + console.error(`query failed: status=${error.status} ${error.message}`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`lease.query(sql, options)` sends the query and resolves with a query handle. +Iterating it with `for await` yields `QwpResultBatch` objects, and the handle's +`completion` promise settles when the query ends: + +| Query option | Default | Purpose | +|---|---|---| +| `binds` | none | Callback that sets the `$1`, `$2`, ... parameters. See [Bind parameters](#bind-parameters). | +| `timeoutMs` | session `queryTimeoutMs` (none) | Deadline that cancels the query. `0` disables it. | +| `initialCredit` | session value (`0`, unbounded) | Flow-control window in bytes. See [Flow control](#flow-control). | +| `autoCredit` | `true` | Replenish the credit window as batches are consumed. | +| `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | + +Iteration and `completion` reject with the same error when the query fails. +Consume the result through `for await`, or `await query.completion` directly +for statements that return no rows. + +A `QwpResultBatch` has: + +- `rowCount` and `columns`: an array of `{ name, type, values, scale?, precisionBits? }`, + where `values` holds one entry per row and `type` is the numeric QWP type + code (compare it with the exported `QWP_COLUMN_TYPE` constants). +- `rows()`: a generator that yields one array of values per row. +- `get(rowIndex, columnIndex)`: one value. + +Batch objects stay valid after iteration moves on, so you can keep them. + +### Reading result values + +Values arrive as these JavaScript types: + +| QuestDB type | JavaScript value | +|---|---| +| BOOLEAN | `boolean` | +| BYTE, SHORT, INT | `number` | +| FLOAT, DOUBLE | `number` | +| LONG | `bigint` | +| TIMESTAMP | `bigint` microseconds since the Unix epoch | +| TIMESTAMP_NS | `bigint` nanoseconds since the Unix epoch | +| DATE | `bigint` milliseconds since the Unix epoch | +| CHAR | one-character `string` | +| VARCHAR, STRING, SYMBOL | `string` | +| BINARY | `Uint8Array` | +| IPv4 | `number`, as a signed 32-bit integer: `192.168.0.1` arrives as `-1062731775`. Use `value >>> 0` for the unsigned address | +| UUID | `{ low: bigint, high: bigint }`, the unsigned low and high 64-bit halves | +| LONG256 | `{ words: [bigint, bigint, bigint, bigint] }`, least significant word first | +| GEOHASH | `{ bits: bigint, precisionBits: number }` | +| DECIMAL | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | +| DOUBLE[], DOUBLE[][], ... | `{ dimensions: number[], values: number[] }` with values in row-major order | +| NULL of any type | `null` | + +INTERVAL values cannot be returned over QWP: the server rejects such a query +with `unsupported column type INTERVAL`. Select the bounds with +`interval_start()` and `interval_end()`, which return timestamps, or cast the +interval with `::varchar`. + +Converting common types: + +```typescript +// TIMESTAMP (bigint microseconds) to Date. Drops sub-millisecond precision. +const toDate = (micros: bigint) => new Date(Number(micros / 1000n)); + +// UUID to its canonical string form +function uuidToString({ low, high }: { low: bigint; high: bigint }): string { + const hex = + high.toString(16).padStart(16, "0") + low.toString(16).padStart(16, "0"); + return [ + hex.slice(0, 8), + hex.slice(8, 12), + hex.slice(12, 16), + hex.slice(16, 20), + hex.slice(20), + ].join("-"); +} + +// IPv4 (signed number) to dotted quad +const ipv4ToString = (value: number) => + [24, 16, 8, 0].map((shift) => ((value >>> 0) >>> shift) & 0xff).join("."); + +console.log(toDate(1723000000000000n).toISOString()); +console.log(uuidToString({ low: 13485158461794337056n, high: 11465204444048149893n })); +console.log(ipv4ToString(-1062731775)); +``` + +### Bind parameters + +Bind values are set by a callback on a `QwpBindValues` object. Indexes are +zero-based (index `0` is `$1`), and setters must be called in ascending index +order without gaps. Every setter returns the object, so calls chain: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol, price FROM trades " + + "WHERE symbol = $1 AND side = $2 AND timestamp >= $3 LIMIT $4", + { + binds: (binds) => + binds + .setVarchar(0, "ETH-USD") + .setVarchar(1, "buy") + .setTimestampMicros(2, BigInt(Date.now() - 3_600_000) * 1000n) + .setLong(3, 1000), + }, + ); + for await (const batch of query) { + for (const row of batch.rows()) console.log(row); + } + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +| Setter | Bind type | +|---|---| +| `setBoolean(index, value)` | BOOLEAN | +| `setByte(index, value)` | BYTE | +| `setShort(index, value)` | SHORT | +| `setChar(index, value)` | CHAR (one-character `string`) | +| `setInt(index, value)` | INT | +| `setLong(index, value)` | LONG (`number` or `bigint`) | +| `setFloat(index, value)` | FLOAT | +| `setDouble(index, value)` | DOUBLE | +| `setDate(index, millis)` | DATE | +| `setTimestampMicros(index, micros)` | TIMESTAMP | +| `setTimestampNanos(index, nanos)` | TIMESTAMP_NS | +| `setVarchar(index, value)` | VARCHAR, STRING, and SYMBOL comparisons. `null` binds NULL | +| `setUuid(index, value)` or `setUuid(index, low, high)` | UUID, as a canonical string or two 64-bit halves. `null` binds NULL | +| `setLong256(index, w0, w1, w2, w3)` | LONG256, least significant word first | +| `setGeohash(index, precisionBits, value)` | GEOHASH | +| `setDecimal64(index, scale, unscaled)` | DECIMAL64 | +| `setDecimal128(index, scale, low, high)` | DECIMAL128 | +| `setDecimal256(index, scale, w0, w1, w2, w3)` | DECIMAL256 | +| `setNull(index, type)` | A typed NULL, with `type` from `QWP_COLUMN_TYPE` | +| `setNullDecimal64/128/256(index, scale)`, `setNullGeohash(index, precisionBits)` | NULL decimals and geohashes, which carry a scale or precision | + +There is no setter for BINARY, IPv4, or arrays. Bind IPv4 as a string and cast +it in SQL (`WHERE ip = $1::ipv4` with `setVarchar`), and pass array values as +SQL literals. + +### DDL and DML statements + +`CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` go through the +same `query()` call. They produce no batches, and `completion` resolves with +`kind: "exec-done"` instead of `kind: "result-end"`: + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const statements = [ + "CREATE TABLE IF NOT EXISTS fills (" + + "timestamp TIMESTAMP, symbol SYMBOL, side SYMBOL, price DOUBLE, amount DOUBLE" + + ") TIMESTAMP(timestamp) PARTITION BY DAY", + "INSERT INTO fills VALUES (now(), 'ETH-USD', 'buy', 2615.54, 0.5)", + "UPDATE fills SET amount = 0.6 WHERE symbol = 'ETH-USD'", + ]; + for (const sql of statements) { + const statement = await lease.query(sql); + const completion = await statement.completion; + if (completion.kind === "exec-done") { + console.log(`${sql.slice(0, 20)}...: ${completion.rowsAffected} rows`); + } + } + } catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + console.error(`statement failed: status=${error.status} ${error.message}`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +| `completion.kind` | Returned for | Fields | +|---|---|---| +| `"result-end"` | Queries that return rows | `totalRows` (`bigint`) | +| `"exec-done"` | DDL and DML | `rowsAffected` (`bigint`, `0` for DDL), `operationType` (QuestDB's numeric statement type) | + +Statements run in order on one lease, because each is awaited before the next +starts, so a `CREATE TABLE` is complete before the `INSERT` that follows it. + +### Cancellation and timeouts -const scale = 4; -const notional = 12345678901234567890n; // represents 1_234_567_890_123_456.7890 +A query ends early in four ways: -await sender - .table("positions") - .symbol("desk", "ny") - .decimalColumnUnscaled("notional", notional, scale) - .atNow(); +- **Deadline.** Set a default with `egressSession: { queryTimeoutMs }`, or per + query with `timeoutMs`. On expiry, iteration and `completion` reject with + `QwpEgressQueryTimeoutError` and the client sends a cancel to QuestDB. +- **Cancel.** `await query.cancel()` asks QuestDB to stop. Iteration and + `completion` reject with `QwpEgressQueryError` whose `status` is `0x0a` + (CANCELLED). +- **Leaving the loop.** `break`, `return`, or an exception inside `for await` + cancels the query, and `completion` rejects with + `QwpEgressQueryAbandonedError`. +- **Waiting without cancelling.** `await query.awaitCompletion(timeoutMs)` + resolves `false` when the wait times out and leaves the query running. + `query.isDone()` reports whether the query has ended. + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryTimeoutError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT symbol, avg(price) FROM trades SAMPLE BY 1m", + { timeoutMs: 5_000 }, + ); + for await (const batch of query) { + console.log(batch.rowCount); + } + } catch (error) { + if (!(error instanceof QwpEgressQueryTimeoutError)) throw error; + console.warn(`query ${error.requestId} timed out after ${error.timeoutMs} ms`); + } finally { + // Waits for the cancellation to drain before the lease is reused. + await lease.close(); + } +} finally { + await db.close(); +} +``` + +After a query ends early, its connection stays busy until QuestDB confirms the +cancellation, and another `query()` on the same lease throws +`a QWP query is already active on this connection`. Return the lease with +`close()` and borrow a new one for the next query. `close()` waits for the +cancellation, up to `query_close_timeout_ms` (5 seconds), and discards the +connection if QuestDB does not confirm in time. + +### Flow control + +By default QuestDB streams results as fast as the network allows, and the +client decodes up to four batches ahead of your loop (`buffer_pool_size`). To +bound how much the server sends ahead, set a byte-credit window with +`initial_credit` in the connect string or `initialCredit` per query: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM trades", { + initialCredit: 1024 * 1024, // server pauses after about 1 MiB + }); + for await (const batch of query) { + // The client replenishes the credit as each batch is consumed. + console.log(batch.rowCount); + } + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +With `autoCredit: false`, call `query.grantCredit(bytes)` yourself. To cap the +rows in each batch, set `max_batch_rows` (1 to 1,048,576). + +### Zero-copy result views + +`query()` materializes every value into JavaScript arrays. For hot paths, +`queryViews()` hands a reusable view of each batch to a callback, and reads +values straight from the received bytes: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + let notional = 0; + const query = await lease.queryViews( + "SELECT timestamp, symbol, price, amount FROM trades", + (batch) => { + const price = batch.column(2); + const amount = batch.column(3); + for (let row = 0; row < batch.rowCount; row++) { + if (!price.isNull(row) && !amount.isNull(row)) { + notional += price.getDouble(row) * amount.getDouble(row); + } + } + // Row-major access reuses one row object for every row. + batch.forEachRow((r) => void r.getSymbol(1)); + }, + ); + await query.completion; + console.log({ notional }); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +The callback is awaited before more credit is granted. The batch, its column +views, and any `Uint8Array` returned from them are valid only until the +callback returns: copy a byte view with `.slice()`, or call +`batch.materialize()`, to keep data. Column views provide typed getters such as +`getBoolean`, `getInt`, `getLong`, `getDouble`, `getString`, `getSymbol`, +`getBinaryView`, and `get` for any type. + +### Compression + +Ask for zstd-compressed results to save bandwidth on large result sets: + +```text +ws::addr=localhost:9000;compression=zstd;compression_level=3; +``` + +`compression` is `raw` (the default), `zstd`, or `auto`; `zstd` and `auto` both +accept a raw reply. `compression_level` ranges from 1 to 22, and the server may +clamp it. `lease.negotiatedCompression` reports what the server chose, for +example `{ codec: "zstd", level: 3 }`. Compression applies to query results +only. + +### Server information + +`lease.serverInfo` describes the server the lease is connected to: `role` (a +`QWP_SERVER_ROLE` value: standalone, primary, replica, or primary catching up), +`zoneId`, `clusterId`, `nodeId`, and `capabilities`. It refreshes after a +failover. + +## Error handling + +### Ingestion errors + +Ingestion reports errors in two ways: + +- **While building a row.** A column method throws, or the promise returned by + `at()` or `atNow()` rejects, with a `TypeError`, `RangeError`, or `Error` for + an invalid value or name. The row in progress is discarded, and the sender + stays usable. +- **Asynchronously, when QuestDB rejects a batch.** The rejection arrives after + `flush()` resolved. It is delivered to the `onSenderError` callback, and + surfaces as a rejection of `waitForAcknowledged()`, or of `flush()` with + `awaitServerAck`. + +```typescript +import { + connectQwpNodeClient, + QWP_SENDER_ERROR_POLICY, + type QwpSenderError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + ingressSession: { + onSenderError: (error: QwpSenderError) => { + const status = error.serverStatusByte?.toString(16); + console.error( + `rejected [${error.category}, policy=${error.appliedPolicy}, ` + + `status=0x${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + + error.serverMessage, + ); + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + // The sender stopped: alert, and fix the data or the schema. + } + }, + }, +}); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + const sequence = await sender.flushAndGetSequence(); + await sender.waitForAcknowledged(sequence, 10_000); + } finally { + // Rethrows a terminal error. The pool then replaces the sender. + await sender.close(); + } +} finally { + await db.close(); +} +``` + +When `onSenderError` is not set, rejections are logged: retriable ones at +`warn`, terminal ones at `error`. Callbacks run asynchronously, never inside the +client's protocol handling, and an exception thrown by a callback is contained. + +`QwpSenderError` fields: + +| Field | Type | Meaning | +|---|---|---| +| `category` | `string` | `schema-mismatch`, `parse-error`, `security-error`, `write-error`, `internal-error`, `not-writable`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `protocol-violation`, `data-loss`, or `unknown`. Branch on this field. | +| `appliedPolicy` | `string` | What the client did: `retriable` (reconnect and resend), `retriable-other` (resend to another endpoint), `terminal` (the sender stopped), or `abandoned` (journaled data was quarantined). | +| `serverStatusByte` | `number` | The raw QWP status code, for example `0x03` for a schema mismatch. Absent for client-side errors. | +| `serverMessage` | `string` | QuestDB's error text, for example `cannot parse DOUBLE from string [value=abc, column=price]`. | +| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `flushAndGetSequence()`. | +| `messageSequence` | `bigint` | The wire sequence of the rejected message. | +| `tableName` | `string` | The table, when the server attributes the rejection to one. Often absent. | +| `detectedAtMs` | `number` | When the client received the rejection. | +| `quarantinedPath` | `string` | For `data-loss` in store-and-forward: where the unreplayable journal was preserved. | + +The default policy follows the category: + +| Category | Policy | Examples | +|---|---|---| +| `schema-mismatch`, `parse-error`, `security-error`, `protocol-violation` | Terminal | Wrong value type for an existing column, malformed data, missing permission | +| `write-error`, `internal-error`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `unknown` | Retriable | Disk pressure, a suspended table, a transient server fault | +| `not-writable` | Retriable on another endpoint | The server is a replica or cannot accept writes | +| `data-loss` | Abandoned | A corrupt store-and-forward journal was set aside | + +A retriable rejection is resent. If the same batch keeps being rejected, after +`max_frame_rejections` (4) attempts spanning at least +`poison_min_escalation_window_millis` (5 minutes), the sender stops as for a +terminal error. The six `on_*_error` connect-string keys are accepted but not +applied by this client. + +**After a terminal error**, the sender is permanently failed. +`waitForAcknowledged()` for the rejected batch rejects with +`QwpIngressNackError`, and every later `flush()` or `close()` rejects with +`QwpReplayRejectedError`, whose `status` and message repeat the server's. Close +the sender and create a new one; the rejected batch is not resent. A pooled +sender is replaced automatically after the `close()` that reports the error. + +Handling notes: + +- **Message stability.** `serverMessage` is free-form English text from the + server. Its wording can change between releases: branch on `category`, not on + the text. +- **Sensitive data.** Server messages can contain column names and values. + Treat them as untrusted input, and redact them before sending them to + third-party error trackers or showing them to end users. +- **Correlation.** There is no server-side request ID. Correlate with the frame + sequence range, `tableName`, and `detectedAtMs`. + +### Query errors + +Query errors reject both the `for await` iteration and `completion`: + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, + QwpEgressQueryTimeoutError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM no_such_table"); + for await (const batch of query) console.log(batch.rowCount); + await query.completion; + } catch (error) { + if (error instanceof QwpEgressQueryError) { + // Prints: 5 [14] table does not exist [table=no_such_table] + console.error(error.status, error.message); + } else if (error instanceof QwpEgressQueryTimeoutError) { + console.error("timed out"); + } else { + throw error; + } + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`QwpEgressQueryError` has `status` (the QWP status code), `message` (the server +text, where parse errors start with the character position in brackets), and +`requestId` (a client-assigned `bigint` that numbers the queries of a +connection). The lease remains usable after a `QwpEgressQueryError`. + +| Status | Name | Meaning | +|---|---|---| +| `0x03` | SCHEMA_MISMATCH | A bind type is incompatible with its placeholder | +| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column | +| `0x06` | INTERNAL_ERROR | Server-side execution failure | +| `0x08` | SECURITY_ERROR | Missing permission | +| `0x0a` | CANCELLED | The query was cancelled with `cancel()` | +| `0x0b` | LIMIT_EXCEEDED | A protocol limit was exceeded | + +Other query errors: + +| Error | Meaning | +|---|---| +| `QwpEgressQueryTimeoutError` | The query deadline expired and cancellation started. Has `requestId` and `timeoutMs`. | +| `QwpEgressQueryAbandonedError` | Iteration ended early, for example with `break`. | +| `QwpEgressQueryCancelTimeoutError` | QuestDB did not confirm a cancellation in time; the connection was closed. | +| `QwpEgressSessionClosedError` | The query connection is closed. | +| `QwpReconnectExhaustedError` | Failover gave up; see [Query failover](#query-failover). | + +As with ingestion, the message text is not stable, may echo parts of the SQL, +and has no server-side correlation ID beyond `requestId`. + +### Connection-level errors + +| Error | Raised when | +|---|---| +| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, `transport`, or `opaque` (browsers, which hide the HTTP status). It also carries `statusCode`, `retryable`, and `url`. | +| `QwpFailoverError` | Every endpoint in a multi-host list failed. `attempts` holds each endpoint and its error. | +| `QwpPoolResourceError` | The pool could not open a new connection. `cause` holds the error above. | +| `QwpPoolAcquireTimeoutError` | Every pooled connection stayed leased past `acquire_timeout_ms`. | +| `QwpReconnectExhaustedError` | The reconnect budget ran out. The sender or query failed permanently. | +| `QwpRoleMismatchError` | No endpoint has the role that `target` requires. | +| `QwpDurableAckUnavailableError` | `request_durable_ack=on`, but the server does not support it. | +| `QwpClientClosedError` | The pooled client, or a returned lease, is already closed. | + +An authentication failure (HTTP 401 or 403) ends the connection attempt for +the whole endpoint list, because a credential rejected by one node is wrong for +all of them, and it is not retried. The exception is a store-and-forward sender +that has connected before: it keeps retrying, so that a rotated credential +cannot strand its journal. Endpoints in error messages have any embedded +credentials removed. + +## Failover and high availability + +:::note Enterprise + +Failing over between several QuestDB hosts requires QuestDB Enterprise +replication. Reconnecting to a single restarted server works in open source +too. + +::: + +### Multiple endpoints + +List several hosts in `addr`: + +```text +wss::addr=db-a.example.com:9000,db-b.example.com:9000,db-c.example.com:9000; +``` -await sender.flush(); +The client ranks endpoints by observed health and by `zone`, and on a +connection loss moves to the next usable one. `addr` is shared by ingestion and +queries. + +Ingestion always needs the primary: replicas refuse writes, and the sender +walks the list until it finds the current primary. Queries can use any node. +`target` selects which roles queries accept (`any`, `primary`, or `replica`), +and `zone` prefers endpoints in the same zone. + +:::caution `target` in the connect string also filters ingestion + +Unlike the Java client, the JavaScript client applies `target` and `zone` from +the connect string to ingestion as well as queries. `target=replica` in the +connect string therefore stops ingestion from reaching the primary. To read +from replicas and write to the primary with one client, keep `target` out of +the connect string and set it for queries only: +`connectQwpNodeClient(conf, { egress: { target: "replica" } })`. + +::: + +### Ingestion reconnect + +When the connection drops, the sender reconnects with exponential backoff and +jitter, then resends every unacknowledged batch: + +| Key | Default | Purpose | +|---|---|---| +| `reconnect_initial_backoff_millis` | `100` | First retry delay. | +| `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | +| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in memory mode. `0` removes the limit. | +| `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | + +Whether the sender gives up depends on the mode (see [Flushing](#flushing)): + +- **Memory mode** retries for up to `reconnect_max_duration_millis` per outage. + When the budget runs out, the sender fails permanently with + `QwpReconnectExhaustedError`, and its unsent rows are lost. +- **Background memory mode** (`initial_connect_retry=async`) and + **store-and-forward** (`sf_dir`) retry indefinitely. + +Setting any `reconnect_*` key also makes the first connection retry within the +budget, as if `initial_connect_retry=on`. Set `initial_connect_retry=off` +explicitly to keep a fail-fast start. + +Replay after a reconnect is at least once: a batch that QuestDB committed just +before the connection dropped is sent again. + + + +### Query failover + +If the connection fails during a query, the client reconnects, to another +endpoint when there is one, and runs the query again from the start: + +| Key | Default | Purpose | +|---|---|---| +| `failover` | `on` | Set `off` to fail the query instead of retrying. | +| `failover_max_attempts` | `8` | Connection attempts per failure. | +| `failover_backoff_initial_ms` | `50` | First retry delay. | +| `failover_backoff_max_ms` | `1000` | Longest delay between retries. | +| `failover_max_duration_ms` | `30000` | Time budget per failure. | + +When the budget runs out, the query rejects with `QwpReconnectExhaustedError`. +A `QwpEgressQueryError` from the server is a query result and never triggers +failover. + +:::warning Clear partial results when a query restarts + +A re-executed query starts again from the first row. Batches that were queued +but not yet consumed are discarded for you, but rows your loop already +processed are delivered again. If your code accumulates rows, register +`onReplayReset` and clear what it collected; otherwise it sees the first part of +the result twice. + +::: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const rows: (readonly unknown[])[] = []; +const db = await connectQwpNodeClient( + "ws::addr=db-a.example.com:9000,db-b.example.com:9000;", + { + egressSession: { + onReplayReset: (event) => { + console.warn(`query ${event.requestId} restarts on ${String(event.endpoint)}`); + rows.length = 0; // discard the partial result + }, + }, + }, +); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM trades LIMIT 100000"); + for await (const batch of query) { + for (const row of batch.rows()) rows.push(row); + } + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +### Connection events + +Register `reconnect.onEvent` to observe connections. Events are delivered +asynchronously through a bounded queue (64 by default, +`connection_listener_inbox_capacity`); when it overflows, the oldest events are +dropped and counted in the metrics. + +```typescript +import { + connectQwpNodeClient, + QWP_RECONNECT_EVENT_KIND, + type QwpReconnectEvent, +} from "@questdb/nodejs-client"; + +function onEvent(event: QwpReconnectEvent) { + switch (event.kind) { + case QWP_RECONNECT_EVENT_KIND.RECONNECTING: + console.warn("connection lost, reconnecting:", event.cause); + break; + case QWP_RECONNECT_EVENT_KIND.FAILED_OVER: + console.warn(`failed over to ${String(event.endpoint)}`); + break; + default: + console.info(event.kind, String(event.endpoint ?? "")); + } +} + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + // Each object replaces the reconnect_* or failover* keys from the connect + // string. Fields left out use the built-in defaults. + ingressSession: { reconnect: { onEvent } }, + egressSession: { reconnect: { onEvent } }, +}); +await db.close(); +``` + +Supplying `egressSession.reconnect`, like setting any `failover*` key, also +makes opening a query connection retry within the failover budget, instead of +failing on the first error. + +| Kind | Meaning | +|---|---| +| `connected` | The first connection succeeded. | +| `reconnecting` | The active connection was lost. `cause` holds the error. | +| `attempt-failed` | One connection attempt failed. The client keeps trying. | +| `reconnected` | Reconnected to the same endpoint. | +| `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | +| `durable-ack-unavailable` | A store-and-forward sender is waiting for an endpoint that supports durable acknowledgement. | +| `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | +| `primary-unavailable` | No reachable endpoint can currently accept writes. | + +`reconnected` and `failed-over` are mutually exclusive: code that tracks the +current node must handle both. Neither `attempt-failed` nor +`primary-unavailable` is terminal: the client keeps retrying until its budget +runs out. + +For ingestion, `ingressSession` also accepts `onProgress`, for published, +acknowledged, and durably acknowledged sequences, and `onError`, for session +errors. `sender.metrics` returns a snapshot of the sender's counters, including +`metrics.ingress` with the replay queue, reconnect, and notification counters. + +## Concurrency + +Node.js runs your code on one thread, but async functions interleave at every +`await`: + +- **`QwpClient`** is safe to share across your whole application. +- **Senders** are not safe for concurrent producers. A row is built across + several calls, so an `await` between `table()` and `at()` lets another task + add columns to the same row. Give each producer its own sender, borrowed from + the pool, and size `sender_pool_max` to match. +- **Query leases** run one query at a time. Borrow one lease per concurrent + query; `query_pool_max` caps concurrent queries. +- **Worker threads** cannot share clients. Create one client per worker, and + give each worker its own `sender_id` when using store-and-forward. + +Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy +work out of them. + +## Browser applications + +`@questdb/browser-client` brings the same ingestion and query API to browsers, +through the browser's native WebSocket. Every sender, query, writer, and +result-batch API above works the same way; the differences are in connecting +and authentication. + +### Connect from the same origin + +QuestDB accepts a browser WebSocket upgrade only when its `Origin` is the same +origin as the request's `Host`, which blocks cross-site WebSocket hijacking. +Serve your application from QuestDB's origin, or route `/write/v4`, `/read/v1`, +and `/exec` to QuestDB through a reverse proxy on your application's origin. +QuestDB answers a cross-origin upgrade with HTTP 400, which the browser reports +as a failed connection. + +If the proxy terminates TLS and forwards plain HTTP to QuestDB, set +[`qwp.browser.tls.termination.enabled`](/docs/configuration/qwp/#browser-connections) +on the server, and forward the browser's `Host` header unchanged. + +Build endpoint URLs from the page location, using `wss:` on HTTPS pages: + +```typescript +const writeUrl = new URL("/write/v4", window.location.href); +writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; +``` + +### Authenticate with a session cookie + +Browsers cannot add an `Authorization` header to a WebSocket upgrade. Instead, +the client authenticates over REST first, and QuestDB sets an HttpOnly +`qdb_session` cookie that the browser sends with the upgrade. When QuestDB has +authentication enabled, put `sessionBootstrap` on the connection options, so +the bootstrap runs before every connection and reconnection attempt: + +```typescript +import { connectQwpBrowserSender } from "@questdb/browser-client"; + +// Obtained by your application, for example from your OIDC provider. +declare const accessToken: string; + +const writeUrl = new URL("/write/v4", window.location.href); +writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; + +const sender = await connectQwpBrowserSender({ + url: writeUrl, + sessionBootstrap: { + // A QuestDB REST token or an OIDC access token. + authentication: { type: "bearer", token: accessToken }, + }, +}); await sender.close(); ``` -`decimalColumnUnscaled` converts `BigInt` inputs into the ILP v3 binary payload. You can also pass an `Int8Array` if you already have a two’s-complement, big-endian byte -array. The scale must stay between 0 and 76, and payloads wider than 32 bytes are rejected up front. This binary path keeps rows compact, making it the preferred option for high-performance feeds. +- Basic authentication uses `{ type: "basic", username, password }`. +- In QuestDB Enterprise, add `serviceAccount: "market_data_writer"` to + `sessionBootstrap` to act as that service account instead of the + authenticated user. +- `bootstrapQwpBrowserSession({ url, authentication })` runs the bootstrap once, + for example right after login. Its `url` is the `/exec` endpoint. +- The bootstrap request uses `credentials: "include"`. The default bootstrap URL + is `/exec` next to the WebSocket endpoint; set `sessionBootstrap.url` when a + proxy exposes it elsewhere. +- A rejected bootstrap fails with `QwpBrowserSessionBootstrapError`, which + carries `statusCode` and `responseBody`. +- The client never reads the HttpOnly cookie. It does not run an OIDC login + flow or refresh tokens: pass a fresh token by creating a new sender or + client. + +### Ingest from a browser -## Configuration options +```typescript +import { connectQwpBrowserSender } from "@questdb/browser-client"; -The minimal configuration string needs to have the protocol, host, and port, as -in: +const writeUrl = new URL("/write/v4", window.location.href); +writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; +const sender = await connectQwpBrowserSender({ url: writeUrl }); +try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.25) + .at(Date.now(), "ms"); + const sequence = await sender.flushAndGetSequence(); + await sender.waitForAcknowledged(sequence, 10_000); +} finally { + await sender.close(); +} ``` -http::addr=localhost:9000; + +`connectQwpBrowserSender(connection, senderOptions, sessionOptions)` takes the +same sender and session options as the typed Node.js API: `autoFlushRows`, +`transactional`, `awaitServerAck`, `onSenderError`, `reconnect`, and so on. +There is no connect string in the browser. + +### Query from a browser + +```typescript +import { + connectQwpBrowserEgress, + QwpEgressQueryError, +} from "@questdb/browser-client"; + +const readUrl = new URL("/read/v1", window.location.href); +readUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; + +const session = await connectQwpBrowserEgress( + { url: readUrl, compression: "zstd" }, + { queryTimeoutMs: 30_000 }, +); +try { + const query = await session.query( + "SELECT timestamp, symbol, price FROM trades WHERE symbol = $1 LIMIT 100", + { + binds: (binds) => binds.setVarchar(0, "ETH-USD"), + // Bound read-ahead: browsers buffer WebSocket frames in memory. + initialCredit: 1024 * 1024, + }, + ); + for await (const batch of query) { + for (const row of batch.rows()) console.log(row); + } + await query.completion; +} catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + console.error(`query failed: status=${error.status} ${error.message}`); +} finally { + await session.close(); +} ``` -For all the extra options you can use, please check -[the client docs](https://questdb.github.io/nodejs-questdb-client/classes/SenderOptions.html) +A session from `connectQwpBrowserEgress()` runs one query at a time, like a +query lease. + +### Pooled browser client -Alternatively, for a breakdown of Configuration string options available across -all clients, see the [Connect string](/docs/connect/clients/connect-string/) page. +`connectQwpBrowserClient()` is the browser counterpart of +`connectQwpNodeClient()`. The `cluster` URL can be an origin, a reverse-proxy +base path, or a `/write/v4` or `/read/v1` URL; the client derives both routes +from it: -## Next Steps +```typescript +import { connectQwpBrowserClient } from "@questdb/browser-client"; + +const clusterUrl = new URL("/", window.location.href); +clusterUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; + +const db = await connectQwpBrowserClient({ + cluster: { url: clusterUrl }, + egress: { compression: "zstd" }, + pool: { senderPoolMax: 2, queryPoolMax: 4 }, +}); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "BTC-USD") + .symbol("side", "sell") + .doubleColumn("price", 39269.98) + .doubleColumn("amount", 0.001) + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } + + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT count() FROM trades"); + for await (const batch of query) console.log(batch.get(0, 0)); + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` -Please refer to the [ILP overview](/docs/connect/compatibility/ilp/overview) for details -about transactions, error control, delivery guarantees, health check, or table -and column auto-creation. +`cluster` also takes `failoverUrls` and `sessionBootstrap`, shared by both +directions. The `ingress` and `egress` sections hold side-specific settings, +such as `requestDurableAck`, `target`, and `compression`. + +### Browser limitations + +| Feature | In browsers | +|---|---| +| Store-and-forward | Not available. Unacknowledged rows are kept in memory and lost when the page closes. | +| UDP | Not available. | +| Connect strings and `QDB_CLIENT_CONF` | Not available. Use the typed options. | +| Custom headers, TLS options | Not available. The browser owns TLS, and the client negotiates through URL parameters and a WebSocket subprotocol instead of headers. | +| Upgrade error details | Not available. A failed upgrade is a `QwpUpgradeError` of kind `opaque`, without an HTTP status. Authentication errors surface from the session bootstrap. | +| `target` and `zone` for ingestion | Not available: browsers cannot see a server's role during the upgrade. They apply to queries only. Do not list replicas for ingestion unless the proxy routes writes to the primary. | +| Durable acknowledgement | Supported through the `requestDurableAck` connection option. | + +## Configuration reference + +The [connect string reference](/docs/connect/clients/connect-string/) documents +every key. The JavaScript client's defaults and deviations: + +| Key | Default | Notes | +|---|---|---| +| `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | +| `username`, `password`, `token` | none | Basic or bearer authentication. | +| `tls_verify`, `tls_roots` | `on`, system store | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | +| `connect_timeout`, `auth_timeout_ms` | `15000` | TCP/TLS connection and upgrade deadlines, in milliseconds. | +| `auto_flush` | `on` | Master switch for the three triggers. | +| `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | +| `auto_flush_interval` | `100` | Milliseconds. `0` disables. `off` is rejected. | +| `auto_flush_bytes` | disabled | Size, or `off`. | +| `close_flush_timeout_millis` | `5000` | ACK wait in a standalone sender's `close()`. | +| `transaction` | `off` | Keep auto-flushed batches in an open transaction until `flush()`. | +| `request_durable_ack` | `off` | Enterprise. | +| `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | +| `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | +| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in memory mode. `0` removes it. | +| `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | +| `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | +| `sf_durability` | `memory` | `memory`, `periodic`, or `append`. | +| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | Journal or memory queue cap. | +| `sf_max_segment_bytes` | `4m` | Journal segment size, which also caps a batch. | +| `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | +| `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. | +| `target`, `zone` | `any`, none | Endpoint role and zone preference. Apply to ingestion too. | +| `failover`, `failover_max_attempts`, `failover_max_duration_ms` | `on`, `8`, `30000` | Query failover. | +| `compression`, `compression_level` | `raw`, `1` | Query result compression. | +| `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | +| `client_id` | `typescript/` | Sent to the server for diagnostics. | +| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client only. | + +The API reference covers every type and option: +[`@questdb/nodejs-client`](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +and +[`@questdb/browser-client`](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_browser-client.html). +The [QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) +in the client repository describes the delivery semantics in depth. + +## Migration + +### From ILP to QWP + +The row API is unchanged, so existing `Sender` code migrates by changing the +connect string and calling `connect()`: + +```diff +- const sender = await Sender.fromConfig("http::addr=localhost:9000"); ++ const sender = await Sender.fromConfig("ws::addr=localhost:9000"); ++ await sender.connect(); +``` -Dive deeper into the Node.js client capabilities, including TypeScript and -Worker Threads examples, by exploring the -[GitHub repository](https://github.com/questdb/nodejs-questdb-client). +| Aspect | ILP over HTTP | QWP over WebSocket | +|---|---|---| +| Connect string schema | `http::`, `https::` | `ws::`, `wss::` | +| Auto-flush rows | 75,000 (600 over TCP) | 1,000 | +| Auto-flush interval | 1,000 ms | 100 ms | +| `flush()` completes when | QuestDB responds to the HTTP request | The batch is published; the ACK arrives later | +| Server rejection | `flush()` throws | Asynchronous: `onSenderError`, `waitForAcknowledged()`, or `flush()` with `awaitServerAck` | +| Rows staged at `close()` | Lost unless flushed | Published, then acknowledged within 5 seconds | +| Reconnect and replay | Retries one request for `retry_timeout` | Automatic, with replay of unacknowledged batches | +| Store-and-forward, querying, pooling | Not available | Available | +| Column types | ILP types | Every QuestDB type | + +Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, +`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss`, +with a hint naming the replacement. To keep ILP-sized batches, set +`auto_flush_rows` and `auto_flush_interval` explicitly. Migrate one sender at a +time: ILP and QWP senders can run side by side. + +### Upgrading from 4.x + +Version 5.0.0 keeps the ILP API and adds QWP. Changes that affect existing ILP +code: + +- **Null values.** Passing `null` or `undefined` to a column or symbol method now + omits the column, which QuestDB stores as NULL. Earlier versions threw a type + error for most such values. Validate data before calling the sender if you + relied on the error. +- **Decimal scale.** `decimalColumn()` over ILP rejects a non-integer `scale` + with a `RangeError`. Earlier versions silently coerced it, writing `2.5` as + scale 2 and `NaN` as scale 0. +- **`intColumn()`** also accepts a `bigint`, for LONG values beyond + `Number.MAX_SAFE_INTEGER`. +- **TCP authentication** now works on Node.js 26, which rejects the JWK the + client previously built. +- **New dependency.** The package now depends on `ws`, used for QWP. + +## ILP transports (legacy) + +The Node.js `Sender` still ingests over ILP, for existing deployments and for +servers without QWP. ILP senders support HTTP (`http::`, `https::`) and TCP +(`tcp::`, `tcps::`) transports: -To learn _The Way_ of QuestDB SQL, see the -[Query & SQL Overview](/docs/query/overview/). +```typescript +import { Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig( + "http::addr=localhost:9000;username=admin;password=quest;", +); +try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.00044) + .at(Date.now(), "ms"); + // ILP does not flush on close: rows still buffered at close() are lost. + await sender.flush(); +} finally { + await sender.close(); +} +``` + +- HTTP connects per request, so `connect()` is not needed; TCP transports + require `await sender.connect()`. `token=...` selects bearer authentication + over HTTP. Over TCP, `username` and `token` set the JWK key ID and private + key. +- `flush()` sends the buffer as one HTTP request, which QuestDB commits as one + transaction, and throws if QuestDB rejects it. +- Decimals need ILP protocol version 3: HTTP negotiates it automatically, and + TCP needs `protocol_version=3`. Arrays need version 2 or later. +- Undici is the default HTTP agent. Set `stdlib_http=on` to use the Node.js + `http` module instead. + +For ILP options, see the +[`SenderOptions` reference](https://questdb.github.io/nodejs-questdb-client/classes/_questdb_nodejs-client.SenderOptions.html) +and the [ILP overview](/docs/connect/compatibility/ilp/overview/). + +## Full example: ingestion and querying with failover + +A production service that ingests trades and queries recent prices, with TLS, +a token, several hosts, error handling, and failover handling: + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, + QwpIngressAckTimeoutError, + QWP_RECONNECT_EVENT_KIND, + type QwpReconnectEvent, + type QwpSenderError, +} from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); + +function logConnection(event: QwpReconnectEvent) { + if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { + console.info("questdb connection:", event.kind, String(event.endpoint ?? "")); + } +} + +const recentPrices: (readonly unknown[])[] = []; + +const db = await connectQwpNodeClient( + "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + + `token=${token};` + + "sender_pool_max=4;query_pool_max=8;", + { + // Queries prefer replicas; ingestion always follows the primary. + egress: { target: "replica", compression: "zstd" }, + ingressSession: { + onSenderError: (error: QwpSenderError) => + console.error("batch rejected:", error.category, error.serverMessage), + // Replaces any reconnect_* keys; omitted fields use the defaults. + reconnect: { onEvent: logConnection }, + }, + egressSession: { + queryTimeoutMs: 30_000, + // Replaces any failover* keys; omitted fields use the defaults. + reconnect: { maxDurationMs: 30_000, onEvent: logConnection }, + onReplayReset: () => { + recentPrices.length = 0; + }, + }, + }, +); + +try { + // Ingestion: one borrowed sender per producer. + const sender = await db.borrowSender(); + try { + for (const [symbol, price, amount] of [ + ["ETH-USD", 2615.54, 0.5], + ["BTC-USD", 39269.98, 0.001], + ] as const) { + await sender + .table("trades") + .symbol("symbol", symbol) + .symbol("side", "buy") + .doubleColumn("price", price) + .doubleColumn("amount", amount) + .at(Date.now(), "ms"); + } + const sequence = await sender.flushAndGetSequence(); + await sender.waitForAcknowledged(sequence, 10_000); + } catch (error) { + if (!(error instanceof QwpIngressAckTimeoutError)) throw error; + console.warn("rows not acknowledged yet; they stay queued for replay"); + } finally { + await sender.close(); + } + + // Querying: rows may not be visible yet, see "Read-after-write". + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol, price FROM trades " + + "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", + { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, + ); + for await (const batch of query) { + for (const row of batch.rows()) recentPrices.push(row); + } + await query.completion; + console.log(recentPrices); + } catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + console.error(`query failed: status=${error.status} ${error.message}`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` -Should you encounter any issues or have questions, the -[Community Forum](https://community.questdb.com/) is a vibrant platform for -discussions. +## Next steps + +- [Connect string reference](/docs/connect/clients/connect-string/) for every + configuration key. +- [Delivery semantics](/docs/concepts/delivery-semantics/) for at-least-once + delivery and deduplication. +- [Client failover](/docs/high-availability/client-failover/concepts/) and + [store-and-forward](/docs/high-availability/store-and-forward/concepts/) + concepts. +- [Query & SQL overview](/docs/query/overview/) for QuestDB SQL. +- The client's + [GitHub repository](https://github.com/questdb/nodejs-questdb-client) and the + [Community Forum](https://community.questdb.com/) for questions and issues. diff --git a/documentation/connect/compatibility/pgwire/nodejs.md b/documentation/connect/compatibility/pgwire/nodejs.md index 5815c4fae2..0a6b68c29a 100644 --- a/documentation/connect/compatibility/pgwire/nodejs.md +++ b/documentation/connect/compatibility/pgwire/nodejs.md @@ -29,10 +29,11 @@ for performance. Our recommendation is to use the `pg` client for most use cases :::tip -For data ingestion, we recommend using QuestDB's first-party clients with -the [InfluxDB Line Protocol (ILP)](/docs/connect/overview/) instead of PGWire. PGWire should primarily be used for -querying data in QuestDB. QuestDB provides an official [JavaScript client](/docs/connect/clients/nodejs/) for data -ingestion using ILP. +For data ingestion, we recommend QuestDB's first-party clients instead of +PGWire. QuestDB provides an official +[JavaScript client](/docs/connect/clients/nodejs/) for Node.js and browsers, +with high-throughput ingestion and streaming SQL queries over QWP. PGWire +remains a good fit when you need a standard PostgreSQL driver or ORM. ::: @@ -781,8 +782,8 @@ QuestDB's support for the PostgreSQL Wire Protocol allows you to use standard Ja time-series data. Both `pg` and `postgres` clients offer good performance and features for working with QuestDB. We recommend the `pg` client for querying. -For data ingestion, consider QuestDB's first-party clients with the InfluxDB Line Protocol (ILP) for maximum -throughput. +For data ingestion, consider the QuestDB [JavaScript client](/docs/connect/clients/nodejs/), which also streams +query results over QWP. Remember that QuestDB is optimized for time-series data, so make the most of its specialized time-series functions like `SAMPLE BY` and `LATEST ON` for efficient queries. diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 2aedc1ee54..2bbcca72ba 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -28,26 +28,25 @@ Pick the path that matches your environment. ## Client Libraries -The first-party libraries for **Java, Python, Go, Rust, Node.js, C & C++, and -.NET** are the recommended way to talk to QuestDB. They speak the -**QuestDB Wire Protocol (QWP)** and unify ingest and query under one -configuration and one connection. +The first-party libraries for **Java, Python, Go, Rust, JavaScript (Node.js +and browsers), C & C++, and .NET** are the recommended way to talk to +QuestDB. They speak the **QuestDB Wire Protocol (QWP)** and unify ingest and +query under one configuration and one connection. ### QWP support -QWP ships in the libraries below. The remaining language clients are being -updated — until they ship a QWP build, they continue to use ILP for ingestion -and PGWire for queries. +QWP ships in every library below. A library marked Beta may still change its +QWP API before it is declared stable. -| Language | QWP support | -| --------- | ----------- | -| Java | ✓ Stable | -| C & C++ | ✓ Stable | -| Rust | ✓ Stable | -| Python | ✓ Stable | -| .NET | Beta | -| Go | Beta | -| Node.js | Planned | +| Language | QWP support | +| --------------------------------- | ----------- | +| Java | ✓ Stable | +| C & C++ | ✓ Stable | +| Rust | ✓ Stable | +| Python | ✓ Stable | +| .NET | Beta | +| Go | Beta | +| JavaScript (Node.js and browsers) | Beta | Highlights: @@ -107,7 +106,8 @@ covering the WebSocket variants for ingress and egress. Read these if you are embedding QuestDB connectivity into an existing framework. QWP also has a UDP transport for fire-and-forget metrics, supported by the -Java, Rust, C and C++ clients via the `udp` connect-string schema. It is +Java, Rust, C and C++ clients, and by the JavaScript client on Node.js, via the +`udp` connect-string schema. It is configured through the [`qwp.udp.*` server settings](/docs/configuration/qwp/#udp-receiver) and is disabled by default; there is no separate byte-level specification page for it. diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index c4e73a3d5e..71d2677fed 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -189,7 +189,7 @@ callers block up to `acquire_timeout_ms` then throw. | `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** | | `reconnect_initial_backoff_millis` | `100` | | `reconnect_max_backoff_millis` | `5000` | -| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python) | +| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/JavaScript) | | `connect_timeout` | unset — per-endpoint TCP connect bound, must be `> 0` | | `auth_timeout_ms` | `15000` | | `max_frame_rejections` | `4` | @@ -312,8 +312,8 @@ here. - Multiple independent senders sharing one `sf_dir` must use distinct `sender_id` values, else the second fails because the slot lock is held. - In pooled `QuestDB` usage, the pool derives per-slot IDs from the base so - pooled senders never collide. The minted name is client-specific: Java uses - `-0`, `-1`, …; the Rust, C and C++ pool uses + pooled senders never collide. The minted name is client-specific: Java and + JavaScript use `-0`, `-1`, …; the Rust, C and C++ pool uses `-ingest-0`, `-ingest-1`, …. - On restart, the cursor engine opens existing segment files and replays unacknowledged frames; acknowledged/truncated frames are not replayed. @@ -374,7 +374,9 @@ has outlasted your configuration. :::note Alignment This is the behaviour of the Java reference client and the .NET client. Other -clients are aligned to it. If you are implementing a new client, the contract +clients are aligned to it, except the JavaScript client in memory mode: a +sender with neither `sf_dir` nor `initial_connect_retry=async` gives up after +`reconnect_max_duration_millis`. If you are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, and apply back-pressure to the producer rather than dropping data. diff --git a/documentation/connect/wire-protocols/qwp-egress-websocket.md b/documentation/connect/wire-protocols/qwp-egress-websocket.md index f981bd4f73..6a91d1a44d 100644 --- a/documentation/connect/wire-protocols/qwp-egress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-egress-websocket.md @@ -33,8 +33,8 @@ For data ingestion, see If your language already has a QuestDB client, use it — the [language client guides](/docs/query/overview) list what's available. The rest of this section is for implementers writing a new one (e.g., to bring -QWP query support to JavaScript, Rust, .NET, or runtimes that the existing -clients don't cover). +QWP query support to Ruby, PHP, or runtimes that the existing clients don't +cover). Compared with the row-oriented HTTP `/exec` JSON endpoint, QWP egress trades a denser binary encoding for higher throughput and lower CPU on both ends: diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index e27c7ea5eb..26edbbaa2c 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -31,8 +31,8 @@ clients, see [QWP egress (WebSocket)](/docs/connect/wire-protocols/qwp-egress-we If your language already has a QuestDB client, use it — the [language client guides](/docs/connect/overview) list what's available. The rest of this section is for implementers writing a new one (e.g., to bring -QWP to JavaScript, Rust, Ruby, .NET, or an embedded runtime that the existing -clients don't cover). +QWP to Ruby, PHP, or an embedded runtime that the existing clients don't +cover). Compared with the line-oriented ILP protocols (`http`, `https`, `tcp`), QWP trades a denser binary encoding for higher throughput and lower CPU on diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 1e708fc691..7180c6bbb3 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -24,8 +24,8 @@ the table below summarises the failover-relevant subset. | Key | Type | Default | Notes | |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | -| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [JavaScript client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. The [JavaScript client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | | `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover the TCP connect or TLS handshake — those use the OS default. Set lower if you have well-known network paths and want faster failover; set higher only if upgrade is genuinely slow. | `addr` syntax — both of these are equivalent and produce the same three-peer diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 6acafca19d..3b8ab82173 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -184,7 +184,7 @@ The exception message distinguishes the two scenarios: `close()` waits up to `close_flush_timeout_millis` for `ackedFsn` to reach `publishedFsn` — i.e. for the server to acknowledge everything the producer has handed in. The default differs by client: 60 s on Java and .NET, 5 s on Rust, -C, C++ and Python. If the wait succeeds, all data is acked. If the timeout +C, C++, Python and JavaScript. If the wait succeeds, all data is acked. If the timeout fires, a `WARN` is logged and: - in **SF mode**, the un-acked tail is left on disk and recovered by the diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 6d11a13303..704873eae3 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -48,11 +48,11 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a JavaScript sender without `sf_dir` or `initial_connect_retry=async` applies it to every outage; see [the JavaScript client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | -| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++ and Python | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | +| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python and JavaScript | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | Cross-reference: [connect-string #reconnect-keys](/docs/connect/clients/connect-string#reconnect-keys). diff --git a/documentation/sidebars.js b/documentation/sidebars.js index e824810980..88f74d1940 100644 --- a/documentation/sidebars.js +++ b/documentation/sidebars.js @@ -76,7 +76,7 @@ module.exports = { { id: "connect/clients/nodejs", type: "doc", - label: "Node.js", + label: "JavaScript", }, { id: "connect/clients/c-and-cpp", diff --git a/shared/clients.json b/shared/clients.json index 649cbf198f..2cabf4b666 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -47,6 +47,14 @@ "logo": "/images/logos/go.svg", "protocol": "QWP" }, + { + "href": "/docs/connect/clients/nodejs", + "name": "JavaScript", + "description": + "QWP ingestion and streaming SQL for Node.js and browsers. Beta.", + "logo": "/images/logos/nodejs-light.svg", + "protocol": "QWP" + }, { "href": "/docs/connect/clients/c-and-cpp", "name": "C & C++", @@ -105,9 +113,9 @@ }, { "href": "/docs/connect/clients/nodejs", - "name": "Node.js", + "name": "JavaScript", "description": - "JavaScript runtime client. Ingests over ILP; QWP support is planned.", + "Node.js client for ILP ingestion over HTTP and TCP, alongside QWP.", "logo": "/images/logos/nodejs-light.svg", "protocol": "ILP" }, From e5c2f6a1bcce93404efe206e639f206f6eca1128 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Mon, 28 Sep 2026 23:33:29 +0100 Subject: [PATCH 02/25] docs(clients): mark JavaScript QWP support as stable The 5.0.0 release ships QWP support as stable. Drop the beta admonition, list the client as stable on the Connect overview and client cards, and state the 5.0.0 minimum version under Requirements. --- documentation/connect/clients/nodejs.md | 12 ++---------- documentation/connect/overview.md | 2 +- shared/clients.json | 16 ++++++++-------- 3 files changed, 11 insertions(+), 19 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index ee12a7e2c6..939983259b 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -39,16 +39,6 @@ Key capabilities: - **Store-and-forward** (Node.js): a disk journal that keeps accepting rows while QuestDB is unreachable and survives process restarts. -:::caution Beta - -QWP support is new in `@questdb/nodejs-client` 5.0.0 and in the first -`@questdb/browser-client` release, and is in beta. Expect the QWP API to change -before it is declared stable, and track the -[client releases](https://github.com/questdb/nodejs-questdb-client/releases). -The ILP transports (`http::`, `https::`, `tcp::`, `tcps::`) are unaffected. - -::: - :::tip Legacy transports The Node.js `Sender` class still speaks ILP over HTTP and TCP. This page @@ -59,6 +49,8 @@ documents the recommended QWP path. For ILP, see ## Requirements +- **`@questdb/nodejs-client` 5.0.0 or newer** for QWP. Earlier versions + support ILP only. - **Node.js 20.18.1 or newer** for `@questdb/nodejs-client`. - **A browser with `WebSocket`, `fetch`, `URL`, `TextEncoder`, and `TextDecoder`** for `@questdb/browser-client`. diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 2bbcca72ba..4cb6ca2e03 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -44,9 +44,9 @@ QWP API before it is declared stable. | C & C++ | ✓ Stable | | Rust | ✓ Stable | | Python | ✓ Stable | +| JavaScript (Node.js and browsers) | ✓ Stable | | .NET | Beta | | Go | Beta | -| JavaScript (Node.js and browsers) | Beta | Highlights: diff --git a/shared/clients.json b/shared/clients.json index 2cabf4b666..c7ec4d8546 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -31,6 +31,14 @@ "logo": "/images/logos/python.svg", "protocol": "QWP" }, + { + "href": "/docs/connect/clients/nodejs", + "name": "JavaScript", + "description": + "QWP ingestion and streaming SQL for Node.js and browsers.", + "logo": "/images/logos/nodejs-light.svg", + "protocol": "QWP" + }, { "href": "/docs/connect/clients/dotnet", "name": ".NET", @@ -47,14 +55,6 @@ "logo": "/images/logos/go.svg", "protocol": "QWP" }, - { - "href": "/docs/connect/clients/nodejs", - "name": "JavaScript", - "description": - "QWP ingestion and streaming SQL for Node.js and browsers. Beta.", - "logo": "/images/logos/nodejs-light.svg", - "protocol": "QWP" - }, { "href": "/docs/connect/clients/c-and-cpp", "name": "C & C++", From 5816730bc90ca1d5b5d925a0d753f18e8b7a7406 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Tue, 29 Sep 2026 10:56:43 +0100 Subject: [PATCH 03/25] docs(clients): correct JavaScript client behavior found in review - IPv4: 0.0.0.0 is rejected, not stored as NULL; pass null instead. - Store-and-forward example: use lazy_connect=on, because the pooled client cannot start while QuestDB is down with initial_connect_retry=async alone. - Terminal rejections: under store-and-forward the rejected batch stays in the journal and every new sender on the slot fails again; document recovery. - Transactions: a borrowed sender's close() commits the open transaction; only a standalone sender rolls back. Clarify that flush() ends the transaction. - target=replica is a strict filter that also fails startup without a replica. - Document QwpSenderCloseTimeoutError from a standalone sender's close(). - TLS verifies against Node's bundled CAs, not the operating system store. - Store-and-forward lock recovery after a crash, including containers. - Query failover: detect re-execution with batchSequence, note that onReplayReset cannot identify the lease and that timeoutMs spans failover. - Add a DEDUP UPSERT KEYS example for at-least-once replay. --- documentation/connect/clients/nodejs.md | 220 ++++++++++++++++++------ 1 file changed, 171 insertions(+), 49 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 939983259b..1d45375c57 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -165,9 +165,10 @@ microseconds since the Unix epoch; see ### Read-after-write -**A flush is not a commit, and a commit is not visibility.** QuestDB -acknowledges ingested rows once they are committed to its write-ahead log, and -applies them to the table asynchronously. A query that runs right after +When `flush()` resolves, the client has published the rows, but QuestDB may not +have received them yet. QuestDB acknowledges a batch once it has committed it +to its write-ahead log, and applies committed rows to the table +asynchronously. A query that runs right after ingestion can therefore fail with `table does not exist` on a first run, or succeed and return no rows. When your code must read its own writes, poll until the rows appear, bounded by a deadline: @@ -284,6 +285,7 @@ try { await sender.flush(); } finally { // Publishes completed rows and waits up to 5 seconds for their ACK. + // Rejects with QwpSenderCloseTimeoutError if the ACK does not arrive. await sender.close(); } ``` @@ -458,11 +460,14 @@ username cannot contain `:`. ### TLS The `wss` schema enables TLS and verifies the server certificate against the -system trust store. Two keys adjust verification, and both are rejected on a -plain `ws` string: +CA certificates bundled with Node.js, not the operating system's trust store. A +private CA installed only in the operating system is not trusted. To trust it, +set `tls_roots`, or add it for the whole process with the +`NODE_EXTRA_CA_CERTS` environment variable, which Node.js reads at startup. +Two keys adjust verification, and both are rejected on a plain `ws` string: - `tls_roots=/path/to/ca.pem` trusts the CA certificates in a PEM file instead - of the system store. The Node.js client accepts PEM only: + of the bundled ones. The Node.js client accepts PEM only: `tls_roots_password` and PKCS#12 or JKS stores are rejected. Export the CA certificates to PEM first. - `tls_verify=unsafe_off` disables certificate verification. Use it only in @@ -503,8 +508,8 @@ const db = await connectQwpNodeClient( `token=${token};` + "tls_roots=/etc/ssl/questdb-ca.pem;", { - // Serve queries from replicas. Keep this out of the connect string: there, - // `target` also applies to ingestion, which must reach the primary. + // Queries run on replicas only (see "Multiple endpoints"). Set target + // here: in the connect string it also applies to ingestion. egress: { target: "replica" }, }, ); @@ -640,8 +645,9 @@ When creating a new pooled connection fails, the borrow rejects with `connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set `lazy_connect=on` to start regardless: senders connect in the background and -buffer rows in memory until QuestDB is reachable, and the query pool stays -empty until the first query. +buffer rows until QuestDB is reachable, in memory or, with `sf_dir`, in the +[store-and-forward](#store-and-forward) journal. The query pool stays empty +until the first query. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -667,8 +673,10 @@ try { ``` `lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, -and rejects an explicit conflicting value. A query borrowed while QuestDB is -still down rejects with `QwpPoolResourceError`. +and rejects an explicit conflicting value. Setting `initial_connect_retry=async` +without `lazy_connect` is not enough: the query pool still connects at startup, +so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query +borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. ## Data ingestion @@ -752,7 +760,7 @@ listed column type when the column does not exist yet: | `binaryColumn(name, value)` | BINARY | `Uint8Array`, copied when staged | | `uuidColumn(name, value)` | UUID | Canonical UUID `string`, or 16 bytes in canonical big-endian order | | `long256Column(name, w0, w1, w2, w3)` | LONG256 | Four 64-bit `bigint` words, least significant first | -| `ipv4Column(name, value)` | IPv4 | Dotted-quad `string` or packed 32-bit `number`. `0.0.0.0` stores NULL | +| `ipv4Column(name, value)` | IPv4 | Dotted-quad `string` or packed 32-bit `number`. `0.0.0.0` is QuestDB's IPv4 NULL value and is rejected; pass `null` for NULL | | `geohashColumn(name, bits, precisionBits)` | GEOHASH | Raw bits as `bigint`, precision from 1 to 60 bits | | `decimalColumnText(name, value)` | DECIMAL(76, scale) | Decimal `string` or `number`. The scale comes from the literal | | `decimalColumn(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, or big-endian two's-complement `Int8Array` | @@ -819,7 +827,8 @@ try { column they do not set. - INT, LONG, and DATE reserve their minimum values as NULL: writing `-2147483648` to INT or `-9223372036854775808n` to LONG or DATE stores NULL. - IPv4 treats `0.0.0.0` as NULL. + IPv4 reserves `0.0.0.0` for NULL too, but `ipv4Column()` rejects it with a + `RangeError` and discards the row: pass `null` to store an IPv4 NULL. - A row where every value is nullish is still sent over WebSocket and stored with NULL in every column. To drop such a row instead, call `cancelRow()` before closing it. Over UDP, `atNow()` rejects such a row while the sender @@ -1067,7 +1076,7 @@ Schema fields: | `binary()` | BINARY | `Uint8Array` | | `uuid()` | UUID | Canonical UUID `string`, 16 big-endian bytes, or `{ low, high }` | | `long256()` | LONG256 | Unsigned 256-bit `bigint`, `0x` hex text, four little-endian words, or `{ words }` | -| `ipv4()` | IPv4 | Dotted-quad `string` or packed `number` | +| `ipv4()` | IPv4 | Dotted-quad `string` or packed `number`. `0.0.0.0` is rejected; omit the field for NULL | | `geohash(precisionBits)` | GEOHASH | Raw bits, base-32 text of `precisionBits / 5` characters, or `{ bits, precisionBits }` | | `decimal64(scale)`, `decimal128(scale)`, `decimal256(scale)` | DECIMAL | Unscaled `bigint`, decimal text, `number`, or `{ unscaled, scale }` | | `doubleArray()` | DOUBLE[] | Nested `number` arrays, or `{ dimensions, values }` | @@ -1125,6 +1134,40 @@ discarded with a warning. On a borrowed sender, `close()` flushes and returns the sender to the pool without waiting; see [Borrowing a sender](#borrowing-a-sender). +If the acknowledgement does not arrive in time, `close()` on a standalone +sender rejects with `QwpSenderCloseTimeoutError`. Its `targetSequence` is the +last published sequence and its `acknowledgedSequence` is how far QuestDB +acknowledged. Without `sf_dir`, the unacknowledged rows are lost. With `sf_dir`, +they stay in the journal for the next sender on that directory. A rejection in +`finally` replaces any error the `try` block threw, so catch it there when that +matters: + +```typescript +import { QwpSenderCloseTimeoutError, Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig("ws::addr=localhost:9000;"); +try { + await sender.connect(); + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.5) + .at(Date.now(), "ms"); +} finally { + try { + await sender.close(); + } catch (error) { + if (!(error instanceof QwpSenderCloseTimeoutError)) throw error; + console.warn( + `published through ${error.targetSequence}, ` + + `acknowledged through ${error.acknowledgedSequence}`, + ); + } +} +``` + ### Awaiting acknowledgements QuestDB acknowledges ingested batches asynchronously. Every published frame gets @@ -1188,7 +1231,7 @@ use [store-and-forward](#store-and-forward) to keep them across restarts. ### Transactions -By default each flushed batch is committed on its own. With transactions on, +By default QuestDB commits each batch on its own. With transactions on, auto-flushed batches stay in an open server-side transaction until you commit: ```typescript @@ -1208,7 +1251,8 @@ try { .floatColumn("amount", 0.01) .at(Date.now(), "ms"); } - // Commits every auto-flushed batch plus the staged rows. + // Ends the transaction: QuestDB commits the auto-flushed batches and the + // staged rows together when it processes this final batch. await sender.flush(); } finally { await sender.close(); @@ -1217,9 +1261,17 @@ try { - The transaction is atomic per table. A flush that spans several tables commits each table separately. -- `flush()` commits. Pooled senders also have `commit()`, an alias of - `flush()`. The typed option is `transactional: true`. -- Closing without committing rolls the open transaction back, with a warning. +- `flush()` ends the transaction: it publishes the final batch, and QuestDB + commits the transaction when it processes that batch. Pooled senders also + have `commit()`, an alias of `flush()`. The typed option is + `transactional: true`. +- Closing a standalone sender without calling `flush()` rolls the open + transaction back, with a warning. +- Returning a borrowed sender with `close()` commits instead, because `close()` + flushes before returning the sender to the pool. `reset()` does not prevent + this: it drops only rows staged since the last flush, not the batches already + sent in the transaction. Use a standalone sender when you may need to abandon + a transaction. - QuestDB does not acknowledge the deferred batches until the commit, so `waitForAcknowledged()` for a sequence inside an open transaction waits for the commit. @@ -1237,7 +1289,7 @@ import { connectQwpNodeClient } from "@questdb/nodejs-client"; const db = await connectQwpNodeClient( "ws::addr=localhost:9000;" + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + - "sf_durability=append;initial_connect_retry=async;", + "sf_durability=append;lazy_connect=on;", ); try { const sender = await db.borrowSender(); @@ -1262,7 +1314,8 @@ try { With a journal, the sender keeps accepting rows while QuestDB is unreachable, until `sf_max_total_bytes` (10 GiB) is full. It retries the connection indefinitely once it has connected, and a new sender opened on the same -directory replays what the previous process left behind. +directory replays what the previous process left behind, once it can take over +the directory's lock (see Lock recovery below). - **Layout.** A standalone `Sender` journals into `/`. A pooled client uses one directory per pooled sender: @@ -1278,8 +1331,27 @@ directory replays what the previous process left behind. - **Backpressure.** When the journal is full, publishing waits up to `sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, then rejects with `QwpReplayStoreAppendTimeoutError`. -- **Startup.** `initial_connect_retry=async` lets a sender start while QuestDB - is down. With the default `off`, the first connection must succeed. +- **Startup.** `lazy_connect=on` lets the pooled client start while QuestDB is + down, as in the example above. `initial_connect_retry=async` alone is not + enough for the pooled client, because its query pool still connects at + startup. A standalone `Sender` needs only `initial_connect_retry=async`. With + the default `off`, the first connection must succeed. +- **Lock recovery.** The Node.js client locks a journal directory with a + `.lock.owner` directory inside it, which records the owner's host name and + process ID, instead of an operating-system file lock. After a crash, a new + sender takes over automatically only when the owner ran on the same host and + its process ID is no longer in use. Otherwise the new sender fails with + `QwpReplayStoreLockedError`. This is common in containers: the application + usually runs as process ID 1, which is in use again after a restart, and a + replacement container usually has a different host name. Once you are sure + the previous process has exited, delete `.lock.owner` from the journal + directory (`/`, or `/-` for a pooled + sender) and start the sender again. +- **Rejected batches.** A batch that QuestDB rejects terminally, such as one + with a value of the wrong type for an existing column, stays in the journal. + Every new sender on that directory, including the pool's replacement for a + failed pooled sender, sends it again and fails the same way. See + [Ingestion errors](#ingestion-errors) for recovery. - **Orphans.** With `drain_orphans=on`, a sender also adopts and drains journals left under the same `sf_dir` by processes that crashed, up to `max_background_drainers` (4) at a time. @@ -1289,6 +1361,27 @@ again, so delivery is at least once: +Create the table with deduplication before ingesting. The upsert keys must +include the designated timestamp and identify a row: rows with equal key values +replace each other. + +```questdb-sql +CREATE TABLE trades ( + timestamp TIMESTAMP, + symbol SYMBOL, + side SYMBOL, + price DOUBLE, + amount DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, symbol, side); +``` + +For an existing table, run +`ALTER TABLE trades DEDUP ENABLE UPSERT KEYS(timestamp, symbol, side);`. +Deduplication recognizes a replayed row only when it carries the same +designated timestamp, so pass event timestamps to `at()` instead of using +`atNow()`. See [Deduplication](/docs/concepts/deduplication/) for choosing keys. + :::warning Do not share a journal directory with a Java client The Node.js client locks journal directories with its own lock files, which the @@ -1409,7 +1502,7 @@ Iterating it with `for await` yields `QwpResultBatch` objects, and the handle's | Query option | Default | Purpose | |---|---|---| | `binds` | none | Callback that sets the `$1`, `$2`, ... parameters. See [Bind parameters](#bind-parameters). | -| `timeoutMs` | session `queryTimeoutMs` (none) | Deadline that cancels the query. `0` disables it. | +| `timeoutMs` | session `queryTimeoutMs` (none) | Deadline that cancels the query. It covers the whole query, including a re-execution after failover. `0` disables it. | | `initialCredit` | session value (`0`, unbounded) | Flow-control window in bytes. See [Flow control](#flow-control). | | `autoCredit` | `true` | Replenish the credit window as batches are consumed. | | `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | @@ -1811,6 +1904,9 @@ try { When `onSenderError` is not set, rejections are logged: retriable ones at `warn`, terminal ones at `error`. Callbacks run asynchronously, never inside the client's protocol handling, and an exception thrown by a callback is contained. +A standalone sender's `close()` can also reject, with +`QwpSenderCloseTimeoutError`, when its rows are not acknowledged in time; see +[Flushing](#flushing). `QwpSenderError` fields: @@ -1845,8 +1941,19 @@ applied by this client. `waitForAcknowledged()` for the rejected batch rejects with `QwpIngressNackError`, and every later `flush()` or `close()` rejects with `QwpReplayRejectedError`, whose `status` and message repeat the server's. Close -the sender and create a new one; the rejected batch is not resent. A pooled -sender is replaced automatically after the `close()` that reports the error. +the sender and create a new one. A pooled sender is replaced automatically +after the `close()` that reports the error. What happens to the rejected batch +depends on the mode: + +- **Without store-and-forward**, the failed sender's unacknowledged batches, + including the rejected one, are discarded with it, and the new sender starts + empty. +- **With store-and-forward**, the rejected batch stays in the journal. Every + new sender on that directory, including the pool's replacement sender, sends + it again and fails the same way. Fix the cause so that QuestDB accepts the + batch, for example by adjusting the table schema, or stop the process and + move the journal directory aside. Moving it aside discards every + unacknowledged batch in it, not only the rejected one. Handling notes: @@ -1965,8 +2072,13 @@ queries. Ingestion always needs the primary: replicas refuse writes, and the sender walks the list until it finds the current primary. Queries can use any node. -`target` selects which roles queries accept (`any`, `primary`, or `replica`), -and `zone` prefers endpoints in the same zone. +`target` selects which roles queries accept: `any` (the default), `primary`, or +`replica`. It is a strict filter, not a preference: with `replica`, queries +never fall back to the primary, and they fail when no replica is reachable, +including against a single open source server. Because the pooled client opens +a query connection at startup, `connectQwpNodeClient()` then rejects too, with +`QwpPoolResourceError` caused by `QwpRoleMismatchError`. Set `query_pool_min=0` +to start without a replica. `zone` prefers endpoints in the same zone. :::caution `target` in the connect string also filters ingestion @@ -2029,35 +2141,37 @@ failover. A re-executed query starts again from the first row. Batches that were queued but not yet consumed are discarded for you, but rows your loop already -processed are delivered again. If your code accumulates rows, register -`onReplayReset` and clear what it collected; otherwise it sees the first part of -the result twice. +processed are delivered again. If your code accumulates rows, clear them when +the query restarts; otherwise it sees the first part of the result twice. ::: +Detect the restart inside the loop. Every batch has a `batchSequence` that +starts at `0n`, and a re-executed query starts again at `0n`. The check works +for each query on its own, so it also covers concurrent queries on separate +leases: + ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const rows: (readonly unknown[])[] = []; const db = await connectQwpNodeClient( "ws::addr=db-a.example.com:9000,db-b.example.com:9000;", - { - egressSession: { - onReplayReset: (event) => { - console.warn(`query ${event.requestId} restarts on ${String(event.endpoint)}`); - rows.length = 0; // discard the partial result - }, - }, - }, ); try { const lease = await db.borrowQuery(); try { - const query = await lease.query("SELECT * FROM trades LIMIT 100000"); + const query = await lease.query("SELECT * FROM trades LIMIT 100000", { + // The deadline covers the whole query, including a re-execution. + timeoutMs: 30_000, + }); + const rows: (readonly unknown[])[] = []; for await (const batch of query) { + // Sequence 0 starts the result, both initially and after a failover. + if (batch.batchSequence === 0n) rows.length = 0; for (const row of batch.rows()) rows.push(row); } await query.completion; + console.log(`${rows.length} rows`); } finally { await lease.close(); } @@ -2066,6 +2180,14 @@ try { } ``` +To be notified of a restart, set `egressSession.onReplayReset` in the second +argument of `connectQwpNodeClient()`. It runs before the first replayed batch +is delivered, and its event has `requestId`, `endpoint`, `previousEndpoint`, +`serverInfo`, and `cause`. The `requestId` matches `query.requestId`, but +request IDs are numbered per connection and every lease of a pooled client +shares the callback, so the event cannot tell concurrent queries apart. Use it +for logging, and the sequence check above to reset results. + ### Connection events Register `reconnect.onEvent` to observe connections. Events are delivered @@ -2352,7 +2474,7 @@ every key. The JavaScript client's defaults and deviations: |---|---|---| | `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | | `username`, `password`, `token` | none | Basic or bearer authentication. | -| `tls_verify`, `tls_roots` | `on`, system store | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | +| `tls_verify`, `tls_roots` | `on`, Node.js CA bundle | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | | `connect_timeout`, `auth_timeout_ms` | `15000` | TCP/TLS connection and upgrade deadlines, in milliseconds. | | `auto_flush` | `on` | Master switch for the three triggers. | | `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | @@ -2500,14 +2622,12 @@ function logConnection(event: QwpReconnectEvent) { } } -const recentPrices: (readonly unknown[])[] = []; - const db = await connectQwpNodeClient( "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + `token=${token};` + "sender_pool_max=4;query_pool_max=8;", { - // Queries prefer replicas; ingestion always follows the primary. + // Queries run on replicas only; ingestion always follows the primary. egress: { target: "replica", compression: "zstd" }, ingressSession: { onSenderError: (error: QwpSenderError) => @@ -2519,9 +2639,8 @@ const db = await connectQwpNodeClient( queryTimeoutMs: 30_000, // Replaces any failover* keys; omitted fields use the defaults. reconnect: { maxDurationMs: 30_000, onEvent: logConnection }, - onReplayReset: () => { - recentPrices.length = 0; - }, + onReplayReset: (event) => + console.warn("query restarts on", String(event.endpoint)), }, }, ); @@ -2559,7 +2678,10 @@ try { "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, ); + const recentPrices: (readonly unknown[])[] = []; for await (const batch of query) { + // A failover re-executes the query from sequence 0: drop the partial result. + if (batch.batchSequence === 0n) recentPrices.length = 0; for (const row of batch.rows()) recentPrices.push(row); } await query.completion; From 56ee2b7f40cb6558f30cb6028047a9afc922a74b Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 14:46:58 +0100 Subject: [PATCH 04/25] docs(clients): scope the QWP client docs to Node.js Only @questdb/nodejs-client ships in this release, so document the Node.js client alone and drop every mention of the browser client: the Browser applications section, the @questdb/browser-client package, installation, and requirements, the browser-only opaque QwpUpgradeError kind, and the server-side Browser connections section in configuration/qwp.md with its changelog entry. Rename the page and sidebar entry back to Node.js, and refer to the Node.js client on the connect-string, failover, store-and-forward, client-behavior, PGWire, and date-to-timestamp pages. --- documentation/changelog.mdx | 3 +- documentation/configuration/qwp.md | 27 +- .../connect/clients/connect-string.md | 40 +-- .../clients/date-to-timestamp-conversion.md | 8 +- documentation/connect/clients/nodejs.md | 297 +++--------------- .../connect/compatibility/pgwire/nodejs.md | 8 +- documentation/connect/overview.md | 31 +- .../wire-protocols/qwp-client-behavior.md | 6 +- .../client-failover/configuration.md | 4 +- .../store-and-forward/concepts.md | 2 +- .../store-and-forward/configuration.md | 4 +- documentation/sidebars.js | 2 +- shared/clients.json | 6 +- 13 files changed, 97 insertions(+), 341 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index ddc3db65f7..5e194b431b 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -27,11 +27,10 @@ This page tracks significant updates to the QuestDB documentation. ### Reference - Added [`cairo.sql.subsample.max.rows`](/docs/configuration/cairo-engine/#cairosqlsubsamplemaxrows), the input row limit for the `lttb`, `m4`, `minmax`, `uniform`, and `cadence` methods of `SUBSAMPLE`. It does not apply to `sdt` -- Added [`qwp.browser.tls.termination.enabled`](/docs/configuration/qwp/#qwpbrowsertlsterminationenabled), for browser QWP connections behind a TLS-terminating reverse proxy, and documented the same-origin check QuestDB applies to [browser connections](/docs/configuration/qwp/#browser-connections) ### Updated -- [JavaScript client](/docs/connect/clients/nodejs/) - Rewrote the Node.js client page for QWP support in `@questdb/nodejs-client` 5.0.0 and the new `@questdb/browser-client` package: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, browser session authentication, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) now lists where the JavaScript client's keys and defaults differ +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) now lists where the Node.js client's keys and defaults differ - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/configuration/qwp.md b/documentation/configuration/qwp.md index 03d72df52b..1928bcf1d1 100644 --- a/documentation/configuration/qwp.md +++ b/documentation/configuration/qwp.md @@ -6,8 +6,7 @@ description: QWP is QuestDB's columnar binary protocol for high-throughput data ingestion (`/write/v4`) and streaming query results (`/read/v1`) over WebSocket and UDP. -These properties control protocol limits, query result compression, browser -connections, and the UDP receiver. WebSocket +These properties control protocol limits and the UDP receiver. WebSocket ingestion and egress share the HTTP server's network settings (port, TLS, worker threads); see [HTTP server configuration](/docs/configuration/http-server/) for those. @@ -53,30 +52,6 @@ overriding the level the client requests via `X-QWP-Accept-Encoding`. `0` disables the override and honours the client's request. Any other value must be in the range `1`-`9`; the server refuses to start otherwise. -## Browser connections - -Browser applications, such as those using the -[JavaScript client](/docs/connect/clients/nodejs/#browser-applications), open -QWP WebSockets from a web page. Browsers always send an `Origin` header with -the upgrade, and QuestDB accepts it only when it is same-origin with the -request's `Host`, including the scheme: an `http://` origin is refused over TLS, -and an `https://` origin is refused over plain HTTP. This blocks cross-site -WebSocket hijacking. Upgrades without an `Origin` header, which is how -non-browser clients connect, are unaffected. Browser QWP connections require a -QuestDB release newer than 10.0.1. - -### qwp.browser.tls.termination.enabled - -- **Default**: `false` -- **Reloadable**: no - -Treats plain-HTTP connections as secure for the browser origin check. Enable it -when a reverse proxy terminates TLS in front of QuestDB: an `https://` origin is -then accepted, and an `http://` origin is refused. The proxy must forward the -browser's original `Host` header unchanged, including the port. The setting -affects the origin check only; it does not mark the `qdb_session` cookie as -`Secure`. - ## UDP receiver :::note diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 0bd62b6b08..15f345bc81 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -256,7 +256,7 @@ Selecting the `wss` schema enables TLS. | Rust, C, C++, Python | PEM (default), JKS, PKCS#12 | | .NET | PKCS#12 / PFX | | Go | none — OS trust store only, both keys rejected at parse time | - | JavaScript (Node.js) | PEM only; `tls_roots_password` is rejected at parse time | + | Node.js | PEM only; `tls_roots_password` is rejected at parse time | - `tls_roots_password` — password for the `tls_roots` file. Required only for a JKS or PKCS#12 trust store; PEM needs no password. Setting it without @@ -292,9 +292,9 @@ Existing JKS and PKCS#12 trust stores keep working through The Go client verifies against the operating-system trust store only and **rejects both keys at parse time**; to trust a private CA there, install it in -the host trust store. The JavaScript client on Node.js accepts PEM only and -rejects `tls_roots_password`: export a JKS or PKCS#12 trust store to PEM -first. On Rust, C, C++ and Python, `tls_roots_password` switches +the host trust store. The Node.js client accepts PEM only and rejects +`tls_roots_password`: export a JKS or PKCS#12 trust store to PEM first. On +Rust, C, C++ and Python, `tls_roots_password` switches the file to a Java keystore and is QWP/WebSocket only: other transports keep PEM as the sole format. Check the relevant [client library page](/docs/connect/overview/#client-libraries) for @@ -323,16 +323,16 @@ act independently: whichever threshold trips first sends the batch. buffered rows. - `auto_flush_rows` — flush when the buffered row count reaches this threshold. Set to `off` to disable. Default where supported: `1000`. - The JavaScript client rejects `off` here; set `0` to disable. + The Node.js client rejects `off` here; set `0` to disable. - `auto_flush_interval` — flush when this many milliseconds have elapsed since the first buffered row. The client evaluates the interval on the next `at()` / `flush()` call, not on a wall-clock timer. Set to `off` to - disable. Default where supported: `100` (100 ms). The JavaScript client + disable. Default where supported: `100` (100 ms). The Node.js client rejects `off` here; set `0` to disable. - `auto_flush_bytes` — flush when the encode buffer reaches this byte size. Set to `off` to disable. Accepts [size suffixes](#size-suffixes). **The default differs by client**: Java - and JavaScript ship it **disabled** (`0`), .NET defaults to `8m` (8 MiB), and + and Node.js ship it **disabled** (`0`), .NET defaults to `8m` (8 MiB), and Rust, C and C++ reject the key outright. A Java application that assumes an 8 MiB byte trigger is active will size batches expecting a flush that never fires. When set to a positive value, the @@ -386,7 +386,7 @@ case-insensitive and 1024-based, matching `-Xmx` conventions: | `g` or `gb` | GiB (× 1024³) | `1g`, `10gb` | | `t` or `tb` | TiB (× 1024⁴) | `1t` | -The JavaScript client accepts only the single-letter suffixes `k`, `m`, `g`, +The Node.js client accepts only the single-letter suffixes `k`, `m`, `g`, and `t`, and rejects `kb`, `mb`, `gb`, and `tb`. ## Multi-host failover {#failover-keys} @@ -431,13 +431,13 @@ single-primary cluster: ingress automatically follows the primary across the host list and adapts when the primary moves to another node. Ingress silently accepts these keys and ignores them. -:::caution JavaScript client +:::caution Node.js client -The JavaScript client applies `target` and `zone` to ingress as well. With +The Node.js client applies `target` and `zone` to ingress as well. With `target=replica` in a shared connect string, its senders accept only replicas and cannot ingest. Set the query-side role through the typed `egress` option instead; see the -[JavaScript client page](/docs/connect/clients/nodejs/#multiple-endpoints). +[Node.js client page](/docs/connect/clients/nodejs/#multiple-endpoints). ::: @@ -518,7 +518,7 @@ equivalent — same architecture, no durability across restarts. |---|---| | Java (`QuestDB` facade) | `/-/` | | Rust, C, C++ (`QuestDb` / `questdb::pool` / `questdb_db`) | `/-ingest-/` | - | JavaScript (`connectQwpNodeClient`) | `/-/` | + | Node.js (`connectQwpNodeClient`) | `/-/` | The minted names belong to that pool's namespace, so pools sharing one `sf_dir` need distinct bases; the slot-in-use error covers both cases @@ -533,7 +533,7 @@ equivalent — same architecture, no durability across restarts. is lost on power failure. The .NET client is the exception: it accepts only `memory` and rejects - anything else at parse time. The JavaScript client also accepts `append`, + anything else at parse time. The Node.js client also accepts `append`, which makes every journal append durable before `flush()` resolves. - `sf_sync_interval_millis` — cadence at which `sf_durability=periodic` checkpoints published frames to stable storage. Default: `5000`. Requires @@ -636,7 +636,7 @@ architecture is that a producer survives an arbitrarily long outage. constructor gives up and returns the error. The running loop and the `async` initial connect never consult it. Default: `300000` (5 min). Setting this enables `initial_connect_retry=on` implicitly; see below. - The JavaScript client differs: a sender with neither `sf_dir` nor + The Node.js client differs: a sender with neither `sf_dir` nor `initial_connect_retry=async` applies this budget to every outage, and fails with `QwpReconnectExhaustedError` when it runs out. - `initial_connect_retry` — whether the client retries the initial connect @@ -659,7 +659,7 @@ architecture is that a producer survives an arbitrarily long outage. milliseconds waiting for buffered frames to drain. Set to `0` or `-1` for fast close (skip the drain). **The default differs by client**: `60000` (60 s) on Java and .NET, `5000` (5 s) on Rust, C, C++ and Python, which - share the same Rust core, and on JavaScript. + share the same Rust core, and on Node.js. This is the shutdown data-loss window. Setting it to `0` skips the drain entirely and drops un-ACKed batches on every clean shutdown. @@ -814,7 +814,7 @@ consumed by the application. Every client's parser accepts the six `on_*_error` keys below, but only clients that implement the policy layer act on them. **In the Java reference -client and the JavaScript client they are currently accepted no-ops** — setting +client and the Node.js client they are currently accepted no-ops** — setting `on_write_error=retriable_other` parses cleanly and changes nothing. .NET does implement them, via `SenderErrorPolicy` and `SenderErrorCategory`. The category table and precedence model below describe the target contract. @@ -877,17 +877,17 @@ description and behaviour notes. | `addr` | `host:port[,host:port…]` | required | [Multi-host failover](#failover-keys) | | `auth_timeout_ms` | int (ms) | `15000` | [Authentication](#auth) | | `auto_flush` | enum (`on` / `off`) | `on` (Rust: only `off`) | [Auto-flushing](#auto-flush) | -| `auto_flush_bytes` | size | Java, JavaScript `0` (off) / .NET `8m` (Rust: rejected) | [Auto-flushing](#auto-flush) | +| `auto_flush_bytes` | size | Java, Node.js `0` (off) / .NET `8m` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `auto_flush_interval` | int (ms) / `off` | `100` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `auto_flush_rows` | int / `off` | `1000` (Rust: rejected) | [Auto-flushing](#auto-flush) | | `buffer_pool_size` | int (≥ 1) | `4` | [Query client keys](#egress-keys) | | `catch_up_cap_gap_min_escalation_window_millis` | int (ms) | `300000` (5 min) | [Store-and-forward](#sf-keys) | | `client_id` | string | client-specific | [Query client keys](#egress-keys) | -| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, JavaScript `5000` | [Ingress reconnect](#reconnect-keys) | +| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, Node.js `5000` | [Ingress reconnect](#reconnect-keys) | | `compression` | enum (`raw` / `zstd` / `auto`) | `raw` | [Query client keys](#egress-keys) | | `compression_level` | int (`1`–`22`) | `1` | [Query client keys](#egress-keys) | | `connect_timeout` | int (ms, `> 0`) | unset | [Authentication](#auth) | -| `connection_listener_inbox_capacity` | int (≥ 1) | `64` (Java, JavaScript) · `256` (Go, .NET) · not supported by Rust, C/C++, Python | [Error handling](#error-handling) | +| `connection_listener_inbox_capacity` | int (≥ 1) | `64` (Java, Node.js) · `256` (Go, .NET) · not supported by Rust, C/C++, Python | [Error handling](#error-handling) | | `drain_orphans` | enum (`on` / `off`) | `off` | [Store-and-forward](#sf-keys) | | `durable_ack_keepalive_interval_millis` | int (ms) | `200` | [Durable ACK](#durable-ack) | | `error_inbox_capacity` | int (≥ 16) | `256` | [Error handling](#error-handling) | @@ -930,7 +930,7 @@ description and behaviour notes. | `sender_pool_min` | int | `1` | [Connection pool](#pool-keys) | | `sf_append_deadline_millis` | int (ms) | `30000` (30 s) | [Store-and-forward](#sf-keys) | | `sf_dir` | path | unset (memory mode) | [Store-and-forward](#sf-keys) | -| `sf_durability` | enum (`memory` / `periodic`) | `memory` (.NET: `memory` only; JavaScript also accepts `append`) | [Store-and-forward](#sf-keys) | +| `sf_durability` | enum (`memory` / `periodic`) | `memory` (.NET: `memory` only; Node.js also accepts `append`) | [Store-and-forward](#sf-keys) | | `sf_max_segment_bytes` | size | `4 MiB` | [Store-and-forward](#sf-keys) | | `sf_max_total_bytes` | size | `128 MiB` mem / `10 GiB` SF | [Store-and-forward](#sf-keys) | | `sf_sync_interval_millis` | int (ms) | `5000` | [Store-and-forward](#sf-keys) | diff --git a/documentation/connect/clients/date-to-timestamp-conversion.md b/documentation/connect/clients/date-to-timestamp-conversion.md index 5e1e2d13fc..0970224ed3 100644 --- a/documentation/connect/clients/date-to-timestamp-conversion.md +++ b/documentation/connect/clients/date-to-timestamp-conversion.md @@ -322,9 +322,9 @@ Learn more about the [QuestDB .NET Client](/docs/connect/clients/dotnet/) ## Date to Timestamp in JavaScript/Node.js A JavaScript `Date` stores milliseconds since the Unix epoch. The QuestDB -JavaScript client takes a timestamp as an integer `number` or a `bigint` -together with a unit: `"ms"`, `"us"` (the default), or `"ns"`. A `Date` -therefore needs no arithmetic: pass `getTime()` with the `"ms"` unit. +Node.js client takes a timestamp as an integer `number` or a `bigint` together +with a unit: `"ms"`, `"us"` (the default), or `"ns"`. A `Date` therefore needs +no arithmetic: pass `getTime()` with the `"ms"` unit. ```javascript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -353,7 +353,7 @@ For an explicit microsecond value, convert through `bigint`: `BigInt(tradeDate.getTime()) * 1000n`. Nanosecond timestamps, with the `"ns"` unit, must be a `bigint`. -Learn more about the [QuestDB JavaScript client](/docs/connect/clients/nodejs/) +Learn more about the [QuestDB Node.js Client](/docs/connect/clients/nodejs/) ## Date to Timestamp in Ruby diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 1d45375c57..79a16f5ddb 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -1,28 +1,20 @@ --- slug: /connect/clients/nodejs -title: JavaScript client for QuestDB -sidebar_label: JavaScript +title: Node.js client for QuestDB +sidebar_label: Node.js description: - "QuestDB JavaScript client for Node.js and browsers: high-throughput data - ingestion and streaming SQL queries over QWP, with pooling, failover, and - store-and-forward." + "QuestDB Node.js client for high-throughput data ingestion and streaming SQL + queries over QWP, with pooling, failover, and store-and-forward." --- import SfDedupWarning from "../../partials/_sf-dedup-warning.partial.mdx" -The QuestDB JavaScript client connects Node.js and browser applications to -QuestDB over [QWP](/docs/connect/wire-protocols/qwp-ingress-websocket/), the -QuestDB Wire Protocol: a columnar binary protocol carried over WebSocket. The -same client ingests data at high throughput and runs SQL queries whose results -stream back as typed, column-oriented batches. - -The client ships as two npm packages built from one code base, with the same -API for ingestion and queries: - -| Package | Runtime | Use it for | -|---|---|---| -| `@questdb/nodejs-client` | Node.js | QWP ingestion and queries over WebSocket, QWP ingestion over UDP, persistent store-and-forward, and the legacy ILP transports | -| `@questdb/browser-client` | Browsers | QWP ingestion and queries over the browser's native WebSocket, with session-cookie authentication. No Node.js dependencies | +The QuestDB Node.js client, `@questdb/nodejs-client`, connects Node.js +applications to QuestDB over +[QWP](/docs/connect/wire-protocols/qwp-ingress-websocket/), the QuestDB Wire +Protocol: a columnar binary protocol carried over WebSocket. The same client +ingests data at high throughput and runs SQL queries whose results stream back +as typed, column-oriented batches. Key capabilities: @@ -36,8 +28,10 @@ Key capabilities: (`db.borrowSender()`) and query leases (`db.borrowQuery()`). - **Failover**: multi-host endpoint lists, automatic reconnect, and replay of unacknowledged rows. -- **Store-and-forward** (Node.js): a disk journal that keeps accepting rows - while QuestDB is unreachable and survives process restarts. +- **Store-and-forward**: a disk journal that keeps accepting rows while + QuestDB is unreachable and survives process restarts. +- **UDP**: fire-and-forget ingestion for metrics where occasional loss is + acceptable. :::tip Legacy transports @@ -51,32 +45,20 @@ documents the recommended QWP path. For ILP, see - **`@questdb/nodejs-client` 5.0.0 or newer** for QWP. Earlier versions support ILP only. -- **Node.js 20.18.1 or newer** for `@questdb/nodejs-client`. -- **A browser with `WebSocket`, `fetch`, `URL`, `TextEncoder`, and - `TextDecoder`** for `@questdb/browser-client`. +- **Node.js 20.18.1 or newer**. - **QuestDB 10.0.0 or newer**, which serves QWP on the HTTP port (`9000` by - default) at `/write/v4` for ingestion and `/read/v1` for queries. - [Browser connections](#browser-applications) need a QuestDB release newer - than 10.0.1: earlier servers reject upgrades from browsers. If QuestDB is not - running yet, see the [quick start](/docs/getting-started/quick-start/). + default) at `/write/v4` for ingestion and `/read/v1` for queries. If QuestDB + is not running yet, see the [quick start](/docs/getting-started/quick-start/). ## Installation -Install the Node.js package: - ```shell npm install @questdb/nodejs-client ``` -For browser applications, install the browser package instead: - -```shell -npm install @questdb/browser-client -``` - -Both packages work with `yarn add` and `pnpm add`. Each exports its complete -API from the package root, ships ES module and CommonJS builds, and bundles -TypeScript declarations. There are no other supported import paths. +The package also installs with `yarn add` and `pnpm add`. It exports its +complete API from the package root, ships ES module and CommonJS builds, and +bundles TypeScript declarations. There are no other supported import paths. The examples on this page are TypeScript ES modules with top-level `await`. They also run as plain JavaScript once type annotations are removed. @@ -222,7 +204,7 @@ Do not replace the poll with a fixed sleep: the apply latency varies with load. ## Connecting -There are three ways to create a client in Node.js: +Create a client with one of these entry points: | Entry point | Returns | Use it for | |---|---|---| @@ -367,7 +349,7 @@ A QWP connect string has the form `schema::key=value;key=value;`: `password=p;;ssw;;rd` sets the password to `p;ssw;rd`. The trailing `;` is optional. -The JavaScript parser differs from some other clients in two places: +The Node.js client's parser differs from some other clients in two places: - `auto_flush_rows` and `auto_flush_interval` take `0`, not `off`, to disable a trigger. `auto_flush=off` disables auto-flushing entirely. @@ -740,8 +722,8 @@ drops every row staged since the last flush. ### Column methods These methods are available on pooled senders and on senders from -`connectQwpNodeSender()` and `connectQwpBrowserSender()`. Each creates the -listed column type when the column does not exist yet: +`connectQwpNodeSender()`. Each creates the listed column type when the column +does not exist yet: | Method | QuestDB type created | Accepted values | |---|---|---| @@ -1111,7 +1093,7 @@ What `flush()` waits for depends on the ingestion mode: |---|---|---|---| | Memory (default) | Neither of the others | The batch is written to the WebSocket, or queued for replay | `flush()` and auto-flushing `at()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | | Background memory | `initial_connect_retry=async`, or `lazy_connect=on` on the pooled client | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | -| Store-and-forward (Node.js) | `sf_dir` | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | +| Store-and-forward | `sf_dir` | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | In every mode, `flush()` does not wait for QuestDB to acknowledge the rows, unless you set `awaitServerAck`. Unacknowledged batches are kept and replayed @@ -1279,9 +1261,9 @@ try { ### Store-and-forward In the default memory mode, unacknowledged rows are lost if the process -exits. Setting `sf_dir` (Node.js only) turns on a disk journal instead: every -batch is appended to the journal before it is sent, a background drainer sends -it in order, and acknowledged segments are deleted. +exits. Setting `sf_dir` turns on a disk journal instead: every batch is +appended to the journal before it is sent, a background drainer sends it in +order, and acknowledged segments are deleted. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -1445,11 +1427,11 @@ try { ``` UDP has no authentication, TLS, acknowledgements, transactions, reconnect, or -store-and-forward, and it is not available in browsers. The server's UDP -receiver is disabled by default; enable it with -[`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver). The default port is -`9007`. `max_datagram_size` (1400 bytes by default) must fit your network path; -a row that cannot fit a datagram fails with `QwpUdpDatagramTooLargeError`. +store-and-forward. The server's UDP receiver is disabled by default; enable it +with [`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver). The default +port is `9007`. `max_datagram_size` (1400 bytes by default) must fit your +network path; a row that cannot fit a datagram fails with +`QwpUdpDatagramTooLargeError`. `multicast_ttl` sets the multicast time-to-live. ## Querying @@ -2032,7 +2014,7 @@ and has no server-side correlation ID beyond `requestId`. | Error | Raised when | |---|---| -| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, `transport`, or `opaque` (browsers, which hide the HTTP status). It also carries `statusCode`, `retryable`, and `url`. | +| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, or `transport`. It also carries `statusCode`, `retryable`, and `url`. | | `QwpFailoverError` | Every endpoint in a multi-host list failed. `attempts` holds each endpoint and its error. | | `QwpPoolResourceError` | The pool could not open a new connection. `cause` holds the error above. | | `QwpPoolAcquireTimeoutError` | Every pooled connection stayed leased past `acquire_timeout_ms`. | @@ -2082,7 +2064,7 @@ to start without a replica. `zone` prefers endpoints in the same zone. :::caution `target` in the connect string also filters ingestion -Unlike the Java client, the JavaScript client applies `target` and `zone` from +Unlike the Java client, the Node.js client applies `target` and `zone` from the connect string to ingestion as well as queries. `target=replica` in the connect string therefore stops ingestion from reaching the primary. To read from replicas and write to the primary with one client, keep `target` out of @@ -2267,208 +2249,10 @@ Node.js runs your code on one thread, but async functions interleave at every Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy work out of them. -## Browser applications - -`@questdb/browser-client` brings the same ingestion and query API to browsers, -through the browser's native WebSocket. Every sender, query, writer, and -result-batch API above works the same way; the differences are in connecting -and authentication. - -### Connect from the same origin - -QuestDB accepts a browser WebSocket upgrade only when its `Origin` is the same -origin as the request's `Host`, which blocks cross-site WebSocket hijacking. -Serve your application from QuestDB's origin, or route `/write/v4`, `/read/v1`, -and `/exec` to QuestDB through a reverse proxy on your application's origin. -QuestDB answers a cross-origin upgrade with HTTP 400, which the browser reports -as a failed connection. - -If the proxy terminates TLS and forwards plain HTTP to QuestDB, set -[`qwp.browser.tls.termination.enabled`](/docs/configuration/qwp/#browser-connections) -on the server, and forward the browser's `Host` header unchanged. - -Build endpoint URLs from the page location, using `wss:` on HTTPS pages: - -```typescript -const writeUrl = new URL("/write/v4", window.location.href); -writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; -``` - -### Authenticate with a session cookie - -Browsers cannot add an `Authorization` header to a WebSocket upgrade. Instead, -the client authenticates over REST first, and QuestDB sets an HttpOnly -`qdb_session` cookie that the browser sends with the upgrade. When QuestDB has -authentication enabled, put `sessionBootstrap` on the connection options, so -the bootstrap runs before every connection and reconnection attempt: - -```typescript -import { connectQwpBrowserSender } from "@questdb/browser-client"; - -// Obtained by your application, for example from your OIDC provider. -declare const accessToken: string; - -const writeUrl = new URL("/write/v4", window.location.href); -writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; - -const sender = await connectQwpBrowserSender({ - url: writeUrl, - sessionBootstrap: { - // A QuestDB REST token or an OIDC access token. - authentication: { type: "bearer", token: accessToken }, - }, -}); -await sender.close(); -``` - -- Basic authentication uses `{ type: "basic", username, password }`. -- In QuestDB Enterprise, add `serviceAccount: "market_data_writer"` to - `sessionBootstrap` to act as that service account instead of the - authenticated user. -- `bootstrapQwpBrowserSession({ url, authentication })` runs the bootstrap once, - for example right after login. Its `url` is the `/exec` endpoint. -- The bootstrap request uses `credentials: "include"`. The default bootstrap URL - is `/exec` next to the WebSocket endpoint; set `sessionBootstrap.url` when a - proxy exposes it elsewhere. -- A rejected bootstrap fails with `QwpBrowserSessionBootstrapError`, which - carries `statusCode` and `responseBody`. -- The client never reads the HttpOnly cookie. It does not run an OIDC login - flow or refresh tokens: pass a fresh token by creating a new sender or - client. - -### Ingest from a browser - -```typescript -import { connectQwpBrowserSender } from "@questdb/browser-client"; - -const writeUrl = new URL("/write/v4", window.location.href); -writeUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; - -const sender = await connectQwpBrowserSender({ url: writeUrl }); -try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.25) - .at(Date.now(), "ms"); - const sequence = await sender.flushAndGetSequence(); - await sender.waitForAcknowledged(sequence, 10_000); -} finally { - await sender.close(); -} -``` - -`connectQwpBrowserSender(connection, senderOptions, sessionOptions)` takes the -same sender and session options as the typed Node.js API: `autoFlushRows`, -`transactional`, `awaitServerAck`, `onSenderError`, `reconnect`, and so on. -There is no connect string in the browser. - -### Query from a browser - -```typescript -import { - connectQwpBrowserEgress, - QwpEgressQueryError, -} from "@questdb/browser-client"; - -const readUrl = new URL("/read/v1", window.location.href); -readUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; - -const session = await connectQwpBrowserEgress( - { url: readUrl, compression: "zstd" }, - { queryTimeoutMs: 30_000 }, -); -try { - const query = await session.query( - "SELECT timestamp, symbol, price FROM trades WHERE symbol = $1 LIMIT 100", - { - binds: (binds) => binds.setVarchar(0, "ETH-USD"), - // Bound read-ahead: browsers buffer WebSocket frames in memory. - initialCredit: 1024 * 1024, - }, - ); - for await (const batch of query) { - for (const row of batch.rows()) console.log(row); - } - await query.completion; -} catch (error) { - if (!(error instanceof QwpEgressQueryError)) throw error; - console.error(`query failed: status=${error.status} ${error.message}`); -} finally { - await session.close(); -} -``` - -A session from `connectQwpBrowserEgress()` runs one query at a time, like a -query lease. - -### Pooled browser client - -`connectQwpBrowserClient()` is the browser counterpart of -`connectQwpNodeClient()`. The `cluster` URL can be an origin, a reverse-proxy -base path, or a `/write/v4` or `/read/v1` URL; the client derives both routes -from it: - -```typescript -import { connectQwpBrowserClient } from "@questdb/browser-client"; - -const clusterUrl = new URL("/", window.location.href); -clusterUrl.protocol = window.location.protocol === "https:" ? "wss:" : "ws:"; - -const db = await connectQwpBrowserClient({ - cluster: { url: clusterUrl }, - egress: { compression: "zstd" }, - pool: { senderPoolMax: 2, queryPoolMax: 4 }, -}); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "BTC-USD") - .symbol("side", "sell") - .doubleColumn("price", 39269.98) - .doubleColumn("amount", 0.001) - .at(Date.now(), "ms"); - } finally { - await sender.close(); - } - - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT count() FROM trades"); - for await (const batch of query) console.log(batch.get(0, 0)); - await query.completion; - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -`cluster` also takes `failoverUrls` and `sessionBootstrap`, shared by both -directions. The `ingress` and `egress` sections hold side-specific settings, -such as `requestDurableAck`, `target`, and `compression`. - -### Browser limitations - -| Feature | In browsers | -|---|---| -| Store-and-forward | Not available. Unacknowledged rows are kept in memory and lost when the page closes. | -| UDP | Not available. | -| Connect strings and `QDB_CLIENT_CONF` | Not available. Use the typed options. | -| Custom headers, TLS options | Not available. The browser owns TLS, and the client negotiates through URL parameters and a WebSocket subprotocol instead of headers. | -| Upgrade error details | Not available. A failed upgrade is a `QwpUpgradeError` of kind `opaque`, without an HTTP status. Authentication errors surface from the session bootstrap. | -| `target` and `zone` for ingestion | Not available: browsers cannot see a server's role during the upgrade. They apply to queries only. Do not list replicas for ingestion unless the proxy routes writes to the primary. | -| Durable acknowledgement | Supported through the `requestDurableAck` connection option. | - ## Configuration reference The [connect string reference](/docs/connect/clients/connect-string/) documents -every key. The JavaScript client's defaults and deviations: +every key. The Node.js client's defaults and deviations: | Key | Default | Notes | |---|---|---| @@ -2500,11 +2284,10 @@ every key. The JavaScript client's defaults and deviations: | `client_id` | `typescript/` | Sent to the server for diagnostics. | | Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client only. | -The API reference covers every type and option: -[`@questdb/nodejs-client`](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) -and -[`@questdb/browser-client`](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_browser-client.html). -The [QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) +The +[API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +covers every type and option. The +[QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) in the client repository describes the delivery semantics in depth. ## Migration diff --git a/documentation/connect/compatibility/pgwire/nodejs.md b/documentation/connect/compatibility/pgwire/nodejs.md index 0a6b68c29a..bc9123385e 100644 --- a/documentation/connect/compatibility/pgwire/nodejs.md +++ b/documentation/connect/compatibility/pgwire/nodejs.md @@ -31,9 +31,9 @@ for performance. Our recommendation is to use the `pg` client for most use cases For data ingestion, we recommend QuestDB's first-party clients instead of PGWire. QuestDB provides an official -[JavaScript client](/docs/connect/clients/nodejs/) for Node.js and browsers, -with high-throughput ingestion and streaming SQL queries over QWP. PGWire -remains a good fit when you need a standard PostgreSQL driver or ORM. +[Node.js client](/docs/connect/clients/nodejs/) with high-throughput ingestion +and streaming SQL queries over QWP. PGWire remains a good fit when you need a +standard PostgreSQL driver or ORM. ::: @@ -782,7 +782,7 @@ QuestDB's support for the PostgreSQL Wire Protocol allows you to use standard Ja time-series data. Both `pg` and `postgres` clients offer good performance and features for working with QuestDB. We recommend the `pg` client for querying. -For data ingestion, consider the QuestDB [JavaScript client](/docs/connect/clients/nodejs/), which also streams +For data ingestion, consider the QuestDB [Node.js client](/docs/connect/clients/nodejs/), which also streams query results over QWP. Remember that QuestDB is optimized for time-series data, so make the most of its specialized time-series functions like diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 4cb6ca2e03..941d3df250 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -28,25 +28,25 @@ Pick the path that matches your environment. ## Client Libraries -The first-party libraries for **Java, Python, Go, Rust, JavaScript (Node.js -and browsers), C & C++, and .NET** are the recommended way to talk to -QuestDB. They speak the **QuestDB Wire Protocol (QWP)** and unify ingest and -query under one configuration and one connection. +The first-party libraries for **Java, Python, Go, Rust, Node.js, C & C++, and +.NET** are the recommended way to talk to QuestDB. They speak the +**QuestDB Wire Protocol (QWP)** and unify ingest and query under one +configuration and one connection. ### QWP support QWP ships in every library below. A library marked Beta may still change its QWP API before it is declared stable. -| Language | QWP support | -| --------------------------------- | ----------- | -| Java | ✓ Stable | -| C & C++ | ✓ Stable | -| Rust | ✓ Stable | -| Python | ✓ Stable | -| JavaScript (Node.js and browsers) | ✓ Stable | -| .NET | Beta | -| Go | Beta | +| Language | QWP support | +| --------- | ----------- | +| Java | ✓ Stable | +| C & C++ | ✓ Stable | +| Rust | ✓ Stable | +| Python | ✓ Stable | +| Node.js | ✓ Stable | +| .NET | Beta | +| Go | Beta | Highlights: @@ -106,9 +106,8 @@ covering the WebSocket variants for ingress and egress. Read these if you are embedding QuestDB connectivity into an existing framework. QWP also has a UDP transport for fire-and-forget metrics, supported by the -Java, Rust, C and C++ clients, and by the JavaScript client on Node.js, via the -`udp` connect-string schema. It is -configured through the [`qwp.udp.*` server +Java, Rust, C, C++ and Node.js clients via the `udp` connect-string schema. It +is configured through the [`qwp.udp.*` server settings](/docs/configuration/qwp/#udp-receiver) and is disabled by default; there is no separate byte-level specification page for it. diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index 71d2677fed..05dd9ae724 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -189,7 +189,7 @@ callers block up to `acquire_timeout_ms` then throw. | `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** | | `reconnect_initial_backoff_millis` | `100` | | `reconnect_max_backoff_millis` | `5000` | -| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/JavaScript) | +| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/Node.js) | | `connect_timeout` | unset — per-endpoint TCP connect bound, must be `> 0` | | `auth_timeout_ms` | `15000` | | `max_frame_rejections` | `4` | @@ -313,7 +313,7 @@ here. `sender_id` values, else the second fails because the slot lock is held. - In pooled `QuestDB` usage, the pool derives per-slot IDs from the base so pooled senders never collide. The minted name is client-specific: Java and - JavaScript use `-0`, `-1`, …; the Rust, C and C++ pool uses + Node.js use `-0`, `-1`, …; the Rust, C and C++ pool uses `-ingest-0`, `-ingest-1`, …. - On restart, the cursor engine opens existing segment files and replays unacknowledged frames; acknowledged/truncated frames are not replayed. @@ -374,7 +374,7 @@ has outlasted your configuration. :::note Alignment This is the behaviour of the Java reference client and the .NET client. Other -clients are aligned to it, except the JavaScript client in memory mode: a +clients are aligned to it, except the Node.js client in memory mode: a sender with neither `sf_dir` nor `initial_connect_retry=async` gives up after `reconnect_max_duration_millis`. If you are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 7180c6bbb3..ab669475f7 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -24,8 +24,8 @@ the table below summarises the failover-relevant subset. | Key | Type | Default | Notes | |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | -| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [JavaScript client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. The [JavaScript client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. The [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | | `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover the TCP connect or TLS handshake — those use the OS default. Set lower if you have well-known network paths and want faster failover; set higher only if upgrade is genuinely slow. | `addr` syntax — both of these are equivalent and produce the same three-peer diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 3b8ab82173..948aebf552 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -184,7 +184,7 @@ The exception message distinguishes the two scenarios: `close()` waits up to `close_flush_timeout_millis` for `ackedFsn` to reach `publishedFsn` — i.e. for the server to acknowledge everything the producer has handed in. The default differs by client: 60 s on Java and .NET, 5 s on Rust, -C, C++, Python and JavaScript. If the wait succeeds, all data is acked. If the timeout +C, C++, Python and Node.js. If the wait succeeds, all data is acked. If the timeout fires, a `WARN` is logged and: - in **SF mode**, the un-acked tail is left on disk and recovered by the diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 704873eae3..a3712a102c 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -48,11 +48,11 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a JavaScript sender without `sf_dir` or `initial_connect_retry=async` applies it to every outage; see [the JavaScript client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender without `sf_dir` or `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | -| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python and JavaScript | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | +| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python and Node.js | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | Cross-reference: [connect-string #reconnect-keys](/docs/connect/clients/connect-string#reconnect-keys). diff --git a/documentation/sidebars.js b/documentation/sidebars.js index 88f74d1940..e824810980 100644 --- a/documentation/sidebars.js +++ b/documentation/sidebars.js @@ -76,7 +76,7 @@ module.exports = { { id: "connect/clients/nodejs", type: "doc", - label: "JavaScript", + label: "Node.js", }, { id: "connect/clients/c-and-cpp", diff --git a/shared/clients.json b/shared/clients.json index c7ec4d8546..0bec77694f 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -33,9 +33,9 @@ }, { "href": "/docs/connect/clients/nodejs", - "name": "JavaScript", + "name": "Node.js", "description": - "QWP ingestion and streaming SQL for Node.js and browsers.", + "Pooled QWP ingestion and streaming SQL for Node.js and TypeScript.", "logo": "/images/logos/nodejs-light.svg", "protocol": "QWP" }, @@ -113,7 +113,7 @@ }, { "href": "/docs/connect/clients/nodejs", - "name": "JavaScript", + "name": "Node.js", "description": "Node.js client for ILP ingestion over HTTP and TCP, alongside QWP.", "logo": "/images/logos/nodejs-light.svg", From f2a991975792cd7f583b489c5836db9896892eb4 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 15:06:59 +0100 Subject: [PATCH 05/25] docs(clients): clarify Node.js journal sharing and memory-mode loss Unacknowledged rows in memory mode may already have reached QuestDB, so say they may be lost rather than that they are lost. Replace the Java-only journal warning: Node.js clients can share an sf_dir, but clients in other languages lock journals with operating-system file locks that the Node.js client does not see, so they may use the directory only after every Node.js client on it has stopped. --- documentation/connect/clients/nodejs.md | 33 ++++++++++++++----------- 1 file changed, 19 insertions(+), 14 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 79a16f5ddb..230a016026 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -1119,10 +1119,10 @@ the sender to the pool without waiting; see If the acknowledgement does not arrive in time, `close()` on a standalone sender rejects with `QwpSenderCloseTimeoutError`. Its `targetSequence` is the last published sequence and its `acknowledgedSequence` is how far QuestDB -acknowledged. Without `sf_dir`, the unacknowledged rows are lost. With `sf_dir`, -they stay in the journal for the next sender on that directory. A rejection in -`finally` replaces any error the `try` block threw, so catch it there when that -matters: +acknowledged. Without `sf_dir`, the unacknowledged rows may be lost. With +`sf_dir`, they stay in the journal for the next sender on that directory. A +rejection in `finally` replaces any error the `try` block threw, so catch it +there when that matters: ```typescript import { QwpSenderCloseTimeoutError, Sender } from "@questdb/nodejs-client"; @@ -1208,8 +1208,9 @@ Acknowledgement is not required for delivery: unacknowledged batches are replayed after a reconnect, and a standalone sender waits for them on `close()`. Wait for acknowledgements when your application must know that QuestDB accepted the rows, for example before committing a source offset. If -the process exits before the acknowledgement, rows still in memory are lost; -use [store-and-forward](#store-and-forward) to keep them across restarts. +the process exits before the acknowledgement, rows still in memory may be +lost; use [store-and-forward](#store-and-forward) to keep them across +restarts. ### Transactions @@ -1260,7 +1261,7 @@ try { ### Store-and-forward -In the default memory mode, unacknowledged rows are lost if the process +In the default memory mode, unacknowledged rows may be lost if the process exits. Setting `sf_dir` turns on a disk journal instead: every batch is appended to the journal before it is sent, a background drainer sends it in order, and acknowledged segments are deleted. @@ -1364,13 +1365,17 @@ Deduplication recognizes a replayed row only when it carries the same designated timestamp, so pass event timestamps to `at()` instead of using `atNow()`. See [Deduplication](/docs/concepts/deduplication/) for choosing keys. -:::warning Do not share a journal directory with a Java client - -The Node.js client locks journal directories with its own lock files, which the -Java client does not see. Never point a running Java client and a running -Node.js client at the same directory. The journal format is shared, so a -directory written by one can be opened by the other after the first has -closed it. +:::warning Share a journal directory only among Node.js clients + +Node.js clients can share an `sf_dir`. Their locks keep each journal to one +process at a time: a second process that opens a journal in use fails with +`QwpReplayStoreLockedError`. Clients in other languages, such as Java, use +operating-system file locks instead, and neither kind of client sees the +other's locks. Such a client must not use the directory while any Node.js +client is running on it: either client could open a journal that the other is +writing, or drain it as an orphan, and corrupt it. Stop every Node.js client on +the directory first. The journal format is shared, so the other client can then +open the journals that the Node.js clients left behind. ::: From ee0c3ae4006fb980b31ece3b21575c9525220607 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 15:20:13 +0100 Subject: [PATCH 06/25] docs(clients): explain when queryViews() callbacks throttle the server Callbacks run one at a time, and credit for a batch is granted only after its callback resolves, which throttles the server only when initialCredit sets a credit window. --- documentation/connect/clients/nodejs.md | 10 +++++++--- 1 file changed, 7 insertions(+), 3 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 230a016026..7dcf86b017 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -1802,9 +1802,13 @@ try { } ``` -The callback is awaited before more credit is granted. The batch, its column -views, and any `Uint8Array` returned from them are valid only until the -callback returns: copy a byte view with `.slice()`, or call +Batches are delivered one at a time: when the callback returns a promise, the +client waits for it before delivering the next batch. With a credit window set +(`initialCredit`), it also grants credit for a batch only after its callback +resolves, so a slow callback throttles the server. + +The batch, its column views, and any `Uint8Array` returned from them are valid +only until the callback returns: copy a byte view with `.slice()`, or call `batch.materialize()`, to keep data. Column views provide typed getters such as `getBoolean`, `getInt`, `getLong`, `getDouble`, `getString`, `getSymbol`, `getBinaryView`, and `get` for any type. From f4b8a1c6cf3ca38e0765ea6ea192ac010e879e89 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 17:39:45 +0100 Subject: [PATCH 07/25] docs(clients): address Node.js QWP review findings --- documentation/changelog.mdx | 2 +- .../connect/clients/connect-string.md | 63 ++-- documentation/connect/clients/nodejs.md | 353 ++++++++++++++---- .../wire-protocols/qwp-client-behavior.md | 10 +- .../wire-protocols/qwp-ingress-websocket.md | 5 +- .../client-failover/concepts.md | 9 +- .../client-failover/configuration.md | 5 +- .../store-and-forward/concepts.md | 13 +- .../store-and-forward/configuration.md | 13 +- .../store-and-forward/operating-and-tuning.md | 38 +- .../store-and-forward/when-to-use.md | 4 +- documentation/query/overview.md | 19 +- plugins/raw-markdown/convert-components.js | 45 ++- .../raw-markdown/convert-components.test.js | 41 ++ 14 files changed, 472 insertions(+), 148 deletions(-) create mode 100644 plugins/raw-markdown/convert-components.test.js diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 5e194b431b..711b267b90 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -30,7 +30,7 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) now lists where the Node.js client's keys and defaults differ +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 15f345bc81..be80fa99ac 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -227,7 +227,7 @@ WebSocket upgrade request. and egress. Bounds the TCP connect phase for each endpoint, so a black-holed host in a multi-host `addr` no longer stalls the [endpoint walk](#failover-keys) until the OS connect timeout. Unset by - default. + default in most clients; Node.js defaults to `15000` (15 seconds). **Mutual TLS (mTLS).** Not supported. The client validates the server's certificate against a trust store but cannot present a client certificate; @@ -248,7 +248,8 @@ Selecting the `wss` schema enables TLS. below. - `tls_roots` — path to a file of trusted root certificates, used instead of the system trust store. If omitted, the client uses the system default - trust store. The accepted on-disk formats are client-specific: + trust store, except the Node.js client, which uses the CA certificates + bundled with Node.js. The accepted on-disk formats are client-specific: | Client | Formats accepted at `tls_roots` | |---|---| @@ -325,10 +326,12 @@ act independently: whichever threshold trips first sends the batch. threshold. Set to `off` to disable. Default where supported: `1000`. The Node.js client rejects `off` here; set `0` to disable. - `auto_flush_interval` — flush when this many milliseconds have elapsed - since the first buffered row. The client evaluates the interval on the - next `at()` / `flush()` call, not on a wall-clock timer. Set to `off` to - disable. Default where supported: `100` (100 ms). The Node.js client - rejects `off` here; set `0` to disable. + since the first buffered row in most clients. The client evaluates the + interval on the next `at()` / `flush()` call, not on a wall-clock timer. + Node.js instead measures from its last flush (or sender creation), so the + first row after an idle period may trigger a flush. Set to `off` to disable. + Default where supported: `100` (100 ms). The Node.js client rejects `off` + here; set `0` to disable. - `auto_flush_bytes` — flush when the encode buffer reaches this byte size. Set to `off` to disable. Accepts [size suffixes](#size-suffixes). **The default differs by client**: Java @@ -392,7 +395,7 @@ and `t`, and rejects `kb`, `mb`, `gb`, and `tb`. ## Multi-host failover {#failover-keys} *Applies to: ingress and egress. The [Role filter and zone preference](#role-filter-and-zone-preference) -sub-section is egress only.* +sub-section is egress only, except on the Node.js client.* :::note QuestDB Enterprise @@ -429,7 +432,8 @@ backoff. Both `target` and `zone` apply to **egress only**. QuestDB is currently a single-primary cluster: ingress automatically follows the primary across the host list and adapts when the primary moves to another node. Ingress -silently accepts these keys and ignores them. +silently accepts these keys and ignores them, except on the Node.js client +(see below). :::caution Node.js client @@ -523,18 +527,13 @@ equivalent — same architecture, no durability across restarts. The minted names belong to that pool's namespace, so pools sharing one `sf_dir` need distinct bases; the slot-in-use error covers both cases (another process or pool holds the slot). -- `sf_durability` — disk durability mode. `memory` (the default) and - `periodic` both ship. `periodic` requires `sf_dir` and checkpoints published - frames in the background at `sf_sync_interval_millis`. `flush` and `append` - are reserved: they parse but are rejected at `build()`. - - Reach for `periodic` when you must survive host loss. `memory` mode is - process-crash durable but **not** host-crash durable, because the page cache - is lost on power failure. - - The .NET client is the exception: it accepts only `memory` and rejects - anything else at parse time. The Node.js client also accepts `append`, - which makes every journal append durable before `flush()` resolves. +- `sf_durability` — disk durability mode. `memory` (the default) relies on the + page cache; `periodic` requires `sf_dir` and checkpoints published frames + in the background at `sf_sync_interval_millis`. Reach for `periodic` when + you must survive host loss: `memory` survives a process crash but not a + power failure. Go and .NET accept only `memory`. Node.js also accepts + `append`, which makes each journal append durable before `flush()` resolves. + Other clients reject `append`; `flush` is not supported. - `sf_sync_interval_millis` — cadence at which `sf_durability=periodic` checkpoints published frames to stable storage. Default: `5000`. Requires `sf_durability=periodic`; rejected otherwise. The configured interval is a @@ -572,6 +571,18 @@ exit, SIGKILL, host crash, or reboot — instantiate a new sender with the parallel with the application's new `append()` calls — it does not block the application. +:::caution Node.js client + +The Node.js client does not use an OS lock. It locks a slot with a +`.lock.owner` directory, so a crashed Node.js sender can leave the slot +locked, and a new sender then fails with `QwpReplayStoreLockedError`. Node.js +and other clients do not see each other's locks: never let them use the same +`sf_dir` at the same time. See the +[Node.js client](/docs/connect/clients/nodejs/#store-and-forward) for lock +recovery. + +::: + If `sf_dir` is a relative path, ensure the process resolves it the same way after restart (typically: use an absolute path). @@ -622,7 +633,9 @@ These keys control the cursor-engine reconnect loop used by QWP ingest. SF mode and memory-only mode share the same loop. A **running** sender retries a transport outage indefinitely with capped exponential backoff — there is no wall-clock give-up: the whole point of the buffering -architecture is that a producer survives an arbitrarily long outage. +architecture is that a producer survives an arbitrarily long outage. The +Node.js client is the exception: in memory mode, it gives up after +`reconnect_max_duration_millis`; see below. - `reconnect_initial_backoff_millis` — initial wait between reconnect attempts. Backoff grows exponentially up to `reconnect_max_backoff_millis`. @@ -886,7 +899,7 @@ description and behaviour notes. | `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, Node.js `5000` | [Ingress reconnect](#reconnect-keys) | | `compression` | enum (`raw` / `zstd` / `auto`) | `raw` | [Query client keys](#egress-keys) | | `compression_level` | int (`1`–`22`) | `1` | [Query client keys](#egress-keys) | -| `connect_timeout` | int (ms, `> 0`) | unset | [Authentication](#auth) | +| `connect_timeout` | int (ms, `> 0`) | unset (Node.js: `15000`) | [Authentication](#auth) | | `connection_listener_inbox_capacity` | int (≥ 1) | `64` (Java, Node.js) · `256` (Go, .NET) · not supported by Rust, C/C++, Python | [Error handling](#error-handling) | | `drain_orphans` | enum (`on` / `off`) | `off` | [Store-and-forward](#sf-keys) | | `durable_ack_keepalive_interval_millis` | int (ms) | `200` | [Durable ACK](#durable-ack) | @@ -917,7 +930,7 @@ description and behaviour notes. | `on_write_error` | enum | `retriable` | [Error handling](#error-handling) | | `pass` | string | unset | [Authentication](#auth) (alias of `password`) | | `password` | string | unset | [Authentication](#auth) | -| `poison_min_escalation_window_millis` | int (ms) | `5000` | [Error handling](#error-handling) | +| `poison_min_escalation_window_millis` | int (ms) | `5000` (Node.js: `300000`) | [Error handling](#error-handling) | | `query_close_timeout_ms` | int (ms) | `5000` | [Query client keys](#egress-keys) | | `query_pool_max` | int | `4` | [Connection pool](#pool-keys) | | `query_pool_min` | int | `1` | [Connection pool](#pool-keys) | @@ -930,12 +943,12 @@ description and behaviour notes. | `sender_pool_min` | int | `1` | [Connection pool](#pool-keys) | | `sf_append_deadline_millis` | int (ms) | `30000` (30 s) | [Store-and-forward](#sf-keys) | | `sf_dir` | path | unset (memory mode) | [Store-and-forward](#sf-keys) | -| `sf_durability` | enum (`memory` / `periodic`) | `memory` (.NET: `memory` only; Node.js also accepts `append`) | [Store-and-forward](#sf-keys) | +| `sf_durability` | enum (`memory` / `periodic` / Node.js `append`) | `memory` (Go and .NET: `memory` only) | [Store-and-forward](#sf-keys) | | `sf_max_segment_bytes` | size | `4 MiB` | [Store-and-forward](#sf-keys) | | `sf_max_total_bytes` | size | `128 MiB` mem / `10 GiB` SF | [Store-and-forward](#sf-keys) | | `sf_sync_interval_millis` | int (ms) | `5000` | [Store-and-forward](#sf-keys) | | `target` | enum (`any` / `primary` / `replica`) | `any` | [Multi-host failover](#failover-keys) | -| `tls_roots` | path | system trust store | [TLS](#tls) | +| `tls_roots` | path | system trust store (Node.js: bundled CAs) | [TLS](#tls) | | `tls_roots_password` | string | unset (JKS / PKCS#12 only) | [TLS](#tls) | | `tls_verify` | enum (`on` / `unsafe_off`) | `on` | [TLS](#tls) | | `token` | string | unset | [Authentication](#auth) | diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 7dcf86b017..15588195b7 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -3,8 +3,8 @@ slug: /connect/clients/nodejs title: Node.js client for QuestDB sidebar_label: Node.js description: - "QuestDB Node.js client for high-throughput data ingestion and streaming SQL - queries over QWP, with pooling, failover, and store-and-forward." + "QuestDB TypeScript and JavaScript Node.js client for high-throughput QWP + ingestion and streaming SQL queries, with pooling, failover, and store-and-forward." --- import SfDedupWarning from "../../partials/_sf-dedup-warning.partial.mdx" @@ -137,8 +137,10 @@ What happens: query handle that is an async iterable of result batches. `batch.rows()` yields one array per row. `query.completion` resolves when the server finishes the query. -4. `db.close()` closes every pooled connection. Pooled senders publish any - remaining rows and wait up to five seconds for QuestDB to acknowledge them. +4. `db.close()` closes the pools. Idle senders publish any remaining rows and + wait up to five seconds for QuestDB to acknowledge them. `db.close()` + resolves even when that wait times out; see + [Closing the pooled client](#closing-the-pooled-client). The table was created automatically by the first row, so its designated timestamp column is named `timestamp`. Timestamps come back as `bigint` @@ -240,7 +242,7 @@ The `QwpClient` handle has five members: | `borrowQuery()` | `Promise` | Lease an exclusive query connection. Its `close()` returns it to the pool. | | `connect()` | `Promise` | Open the pool minimums. Called for you by `connectQwpNodeClient()`. Safe to retry after a failure. | | `metrics` | `QwpClientMetrics` | Pool counters (`total`, `available`, `leased`, `creating`, `waiting`) for senders and queries. | -| `close()` | `Promise` | Reject new borrows, close idle connections, cancel active queries, and close the pools. Idempotent. | +| `close()` | `Promise` | Reject new borrows, cancel active queries, close the query connections and idle senders, and wait up to 5 seconds for borrowed senders to be returned. Resolves even if rows are not acknowledged; see [Closing the pooled client](#closing-the-pooled-client). Idempotent. | Share one `QwpClient` across your application and close it at shutdown. See [The connection pool](#the-connection-pool) for pool sizing and lease rules. @@ -463,8 +465,11 @@ combined with `tls_verify` or `tls_roots`. Two deadlines bound connection setup: `connect_timeout` covers DNS and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits -`connect_timeout` when only the latter is set. A timeout fails with -`QwpUpgradeError`, whose `timeoutPhase` is `connect` or `authentication`. +`connect_timeout` when only the latter is set. A timeout produces a +`QwpUpgradeError` whose `timeoutPhase` is `connect` or `authentication`. The +pooled client reports it, like every failure to open a connection, as the +`cause` of a `QwpPoolResourceError`; see +[Connection-level errors](#connection-level-errors). ### Unsupported authentication paths @@ -488,10 +493,13 @@ if (!token) throw new Error("QDB_TOKEN is not set"); const db = await connectQwpNodeClient( "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + `token=${token};` + - "tls_roots=/etc/ssl/questdb-ca.pem;", + "tls_roots=/etc/ssl/questdb-ca.pem;" + + // Start, and ingest, even while no replica is reachable. + "query_pool_min=0;", { - // Queries run on replicas only (see "Multiple endpoints"). Set target - // here: in the connect string it also applies to ingestion. + // Queries run on replicas only, with no fallback to the primary (see + // "Multiple endpoints"). Set target here: in the connect string it also + // applies to ingestion. egress: { target: "replica" }, }, ); @@ -502,6 +510,9 @@ try { } ``` +With `query_pool_min=0`, the client starts while no replica is reachable, and +a query borrowed during that time rejects with `QwpPoolResourceError`. + ## The connection pool The pooled client keeps two elastic pools: one of senders and one of query @@ -550,8 +561,9 @@ that hold a sender at the same time. `close()` on a borrowed sender flushes completed rows, discards an unfinished row with a warning, and returns the sender to the pool. It does not close the WebSocket and does not wait for QuestDB to acknowledge the rows. To confirm -delivery before returning the sender, use -[`flushAndGetSequence()` and `waitForAcknowledged()`](#awaiting-acknowledgements). +delivery before returning the sender, call `flush()` and then +`waitForAcknowledged(sender.publishedSequence)`; see +[Awaiting acknowledgements](#awaiting-acknowledgements). When a borrowed sender's `close()` fails, the pool discards that sender and opens a new one for the next borrow. Because QuestDB reports rejected batches @@ -659,6 +671,31 @@ and rejects an explicit conflicting value. Setting `initial_connect_retry=async` without `lazy_connect` is not enough: the query pool still connects at startup, so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. +Without `sf_dir`, the buffered rows exist only in memory, and they are lost if +the client closes before QuestDB becomes reachable; see +[Closing the pooled client](#closing-the-pooled-client). + +### Closing the pooled client + +`db.close()` rejects new borrows, then: + +- Cancels active queries and closes every query connection, including leased + ones. +- Closes idle senders. Each publishes its remaining rows and waits up to + `close_flush_timeout_millis` (5 seconds) for QuestDB to acknowledge them. +- Waits up to 5 seconds, or `acquire_timeout_ms` if lower, for borrowed + senders to be returned. A sender still borrowed after that stays open, and + its owner must `close()` it. + +`db.close()` resolves even when an acknowledgement does not arrive in time. It +reports the timeout to `ingressSession.onError` as a non-terminal +`QwpIngressAckTimeoutError`, which is logged as a warning by default. Without +`sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they stay in +the journal, and the next sender on the same directory replays them. To know +that QuestDB accepted every row before shutting down, wait for the +acknowledgement before returning each sender (see +[Awaiting acknowledgements](#awaiting-acknowledgements)), or use +[store-and-forward](#store-and-forward). ## Data ingestion @@ -925,6 +962,9 @@ Create decimal columns ahead of time with the precision you need. QWP can create them automatically, but it picks the maximum precision of the wire width (18, 38, or 76 digits). See [decimal data type](/docs/query/datatypes/decimal/#creating-tables-with-decimals). +To also query a decimal column over QWP, give it a precision of 10 or more: +current servers cannot return a DECIMAL with a precision of 9 or less (see +[Reading result values](#reading-result-values)). ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -1106,8 +1146,17 @@ then rejects with `QwpMemoryReplayAppendTimeoutError`. Tune the cap with `sf_max_total_bytes` and the wait with `sf_append_deadline_millis`; without `sf_dir` they size the memory queue. Watch `sender.metrics.ingress` (`memoryReplayUsedBytes`, `totalMemoryReplayBackpressureStalls`) to detect -backpressure before it blocks. A single row larger than the server's batch -limit is rejected before it is sent, with `QwpBatchTooLargeError`. +backpressure before it blocks. + +**Oversized rows.** When the sender connects, QuestDB advertises the largest +batch it accepts: about 2 MiB on a default server, set by +`http.recv.buffer.size`. With `sf_dir`, a batch must also fit in +`sf_max_segment_bytes`. A row too large to fit in one batch fails the +`flush()`, or the `at()` whose auto-flush sends it, with +`QwpBatchTooLargeError` before anything is sent. The staged rows are kept, so +every later flush fails the same way, and `close()` discards them and rejects +with the same error. Call `reset()` to drop every row staged since the last +flush, then write the other rows again. **Closing.** `close()` on a standalone sender publishes completed rows and waits up to `close_flush_timeout_millis` (5 seconds by default) for their @@ -1174,9 +1223,10 @@ try { .doubleColumn("amount", 0.5) .at(Date.now(), "ms"); - const sequence = await sender.flushAndGetSequence(); - // Rejects with the server's error if QuestDB rejected the batch. - await sender.waitForAcknowledged(sequence, 10_000); + // Wait for every row published so far, including rows an auto-flush + // already sent. Rejects with the server's error if QuestDB rejected them. + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } catch (error) { if (error instanceof QwpIngressAckTimeoutError) { // Not acknowledged in time. The rows are still pending, not lost. @@ -1194,10 +1244,22 @@ try { | Member | Returns | |---|---| -| `flushAndGetSequence()` | Publishes staged rows and resolves with the highest sequence (`bigint`) this call published, or `-1n` when there was nothing to publish. | -| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout, without closing the sender, or with the server's rejection. | +| `publishedSequence` | The highest sequence this sender published, including by auto-flushes, or `-1n`. After `flush()`, it covers every row written so far. | +| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout (15 seconds by default), without closing the sender, or with the server's rejection. | | `acknowledgedSequence` | The highest acknowledged sequence, or `-1n`. | -| `publishedSequence` | The highest published sequence, or `-1n`. | +| `flushAndGetSequence()` | Publishes staged rows and resolves with the highest sequence (`bigint`) this call published, or `-1n` when there was nothing to publish. Rows an earlier auto-flush published are not covered. | + +:::caution Do not wait on the result of `flushAndGetSequence()` + +An auto-flush inside `at()` publishes the staged rows on its own: on the row +that reaches `auto_flush_rows`, or on the first row after the sender was idle +for `auto_flush_interval` (100 ms by default). `flushAndGetSequence()` then +has nothing left to publish and returns `-1n`, and `waitForAcknowledged(-1n)` +resolves at once, before QuestDB has acknowledged or rejected the rows. To +wait for every row written so far, call `flush()` and wait for +`publishedSequence`, as in the example above. + +::: To make every `flush()` wait for its acknowledgement, set `awaitServerAck`: `connectQwpNodeClient(conf, { sender: { awaitServerAck: true } })`, or @@ -1244,6 +1306,16 @@ try { - The transaction is atomic per table. A flush that spans several tables commits each table separately. +- A transaction is atomic only up to a size limit. QuestDB commits a table + early once its open transaction holds + [`qwp.max.uncommitted.rows`](/docs/configuration/qwp/#qwpmaxuncommittedrows) + rows (1,000,000 by default), and closing without `flush()` cannot roll back + what it committed. The open transaction's batches also stay in the replay + queue until the commit, so they must fit in `sf_max_total_bytes` (128 MiB + without `sf_dir`). Beyond that, publishing waits `sf_append_deadline_millis` + (30 seconds) and then rejects with `QwpMemoryReplayAppendTimeoutError`, or + `QwpReplayStoreAppendTimeoutError` with `sf_dir`. Split large loads into + several transactions. - `flush()` ends the transaction: it publishes the final batch, and QuestDB commits the transaction when it processes that batch. Pooled senders also have `commit()`, an alias of `flush()`. The typed option is @@ -1266,6 +1338,26 @@ exits. Setting `sf_dir` turns on a disk journal instead: every batch is appended to the journal before it is sent, a background drainer sends it in order, and acknowledged segments are deleted. +Before ingesting, create a deduplicated table while QuestDB is reachable. +Use both the event timestamp and a stable, source-assigned trade ID as upsert +keys: distinct trades can share a millisecond timestamp, symbol, and side. + +```questdb-sql +CREATE TABLE trades_sf ( + timestamp TIMESTAMP, + trade_id SYMBOL, + symbol SYMBOL, + side SYMBOL, + price DOUBLE, + amount DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, trade_id); +``` + +Pass the same source ID and timestamp again if the application retries an +event. The following values represent one source event; do not regenerate them +when retrying it: + ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -1274,16 +1366,18 @@ const db = await connectQwpNodeClient( "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + "sf_durability=append;lazy_connect=on;", ); +const event = { tradeId: "trade-12345", timestampMs: 1723000000000 }; try { const sender = await db.borrowSender(); try { await sender - .table("trades") + .table("trades_sf") + .symbol("trade_id", event.tradeId) .symbol("symbol", "ETH-USD") .symbol("side", "buy") .doubleColumn("price", 2615.54) .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); + .at(event.timestampMs, "ms"); // Resolves once the rows are in the journal, even if QuestDB is down. await sender.flush(); } finally { @@ -1344,26 +1438,14 @@ again, so delivery is at least once: -Create the table with deduplication before ingesting. The upsert keys must -include the designated timestamp and identify a row: rows with equal key values -replace each other. - -```questdb-sql -CREATE TABLE trades ( - timestamp TIMESTAMP, - symbol SYMBOL, - side SYMBOL, - price DOUBLE, - amount DOUBLE -) TIMESTAMP(timestamp) PARTITION BY DAY -DEDUP UPSERT KEYS(timestamp, symbol, side); -``` - -For an existing table, run -`ALTER TABLE trades DEDUP ENABLE UPSERT KEYS(timestamp, symbol, side);`. +The `trades_sf` keys identify a trade without collapsing distinct trades +that share a millisecond timestamp, symbol, and side. For an existing table +that already has a stable `trade_id` column, enable deduplication with +`ALTER TABLE trades_sf DEDUP ENABLE UPSERT KEYS(timestamp, trade_id);`. Deduplication recognizes a replayed row only when it carries the same -designated timestamp, so pass event timestamps to `at()` instead of using -`atNow()`. See [Deduplication](/docs/concepts/deduplication/) for choosing keys. +designated timestamp and trade ID, so reuse event values on application retries +instead of calling `atNow()` or generating a new ID. See +[Deduplication](/docs/concepts/deduplication/) for choosing keys. :::warning Share a journal directory only among Node.js clients @@ -1403,7 +1485,11 @@ wss::addr=db.example.com:9000;token=YOUR_TOKEN;request_durable_ack=on; To make every flush wait for durability, add the typed option `{ sender: { awaitDurableAck: true } }`. If the server does not support durable -acknowledgement, connecting fails with `QwpDurableAckUnavailableError`. +acknowledgement, connecting fails with `QwpDurableAckUnavailableError`, which +the pooled client reports as the `cause` of a `QwpPoolResourceError`. A sender +that connects in the background, with `initial_connect_retry=async` or +`lazy_connect=on`, does not fail: it keeps retrying and emits +`durable-ack-unavailable` [connection events](#connection-events). ### Fire-and-forget UDP @@ -1435,8 +1521,10 @@ UDP has no authentication, TLS, acknowledgements, transactions, reconnect, or store-and-forward. The server's UDP receiver is disabled by default; enable it with [`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver). The default port is `9007`. `max_datagram_size` (1400 bytes by default) must fit your -network path; a row that cannot fit a datagram fails with -`QwpUdpDatagramTooLargeError`. +network path. A row that cannot fit in a datagram fails the flush with +`QwpUdpDatagramTooLargeError`. As with an +[oversized WebSocket batch](#flushing), the staged rows are kept, so later +flushes and `close()` fail too: call `reset()` to drop them. `multicast_ttl` sets the multicast time-to-live. ## Querying @@ -1528,14 +1616,21 @@ Values arrive as these JavaScript types: | UUID | `{ low: bigint, high: bigint }`, the unsigned low and high 64-bit halves | | LONG256 | `{ words: [bigint, bigint, bigint, bigint] }`, least significant word first | | GEOHASH | `{ bits: bigint, precisionBits: number }` | -| DECIMAL | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | +| DECIMAL with a precision of 10 or more | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | | DOUBLE[], DOUBLE[][], ... | `{ dimensions: number[], values: number[] }` with values in row-major order | | NULL of any type | `null` | -INTERVAL values cannot be returned over QWP: the server rejects such a query -with `unsupported column type INTERVAL`. Select the bounds with -`interval_start()` and `interval_end()`, which return timestamps, or cast the -interval with `::varchar`. +Some column types cannot be returned over QWP. The server rejects such a query +with status `0x06` and a message such as `unsupported column type INTERVAL`. +Convert the column in SQL instead: + +- INTERVAL: select the bounds with `interval_start()` and `interval_end()`, + which return timestamps, or cast the interval with `::varchar`. +- DECIMAL with a precision of 9 or less, which QuestDB stores as DECIMAL8, + DECIMAL16, or DECIMAL32: cast it to a wider precision, for example + `price::DECIMAL(18, 2)`. +- An untyped `NULL` literal, as in `SELECT NULL`: give it a type, for example + `NULL::double`. Converting common types: @@ -1622,7 +1717,7 @@ try { | `setDecimal64(index, scale, unscaled)` | DECIMAL64 | | `setDecimal128(index, scale, low, high)` | DECIMAL128 | | `setDecimal256(index, scale, w0, w1, w2, w3)` | DECIMAL256 | -| `setNull(index, type)` | A typed NULL, with `type` from `QWP_COLUMN_TYPE` | +| `setNull(index, type)` | A typed NULL. `type` is one of the scalar `QwpBindType` values in `QWP_COLUMN_TYPE`, not every column type: BINARY, IPv4, arrays, STRING, and SYMBOL are excluded (bind text as VARCHAR). | | `setNullDecimal64/128/256(index, scale)`, `setNullGeohash(index, precisionBits)` | NULL decimals and geohashes, which carry a scale or precision | There is no setter for BINARY, IPv4, or arrays. Bind IPv4 as a string and cast @@ -1633,7 +1728,19 @@ SQL literals. `CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` go through the same `query()` call. They produce no batches, and `completion` resolves with -`kind: "exec-done"` instead of `kind: "result-end"`: +`kind: "exec-done"` instead of `kind: "result-end"`. + +:::warning DDL and DML can run twice with query failover + +With the default `failover=on`, a connection loss replays any in-flight SQL, +including DDL and DML. QuestDB may have applied an `INSERT` before its +`exec-done` response was lost, so replay can insert it again. A transport +error does not prove the statement failed. Use a separate client with +`failover=off` for non-idempotent statements, as below, and verify an +uncertain outcome before retrying manually. Alternatively, make the SQL +idempotent; see [Query failover](#query-failover). + +::: ```typescript import { @@ -1641,7 +1748,7 @@ import { QwpEgressQueryError, } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); try { const lease = await db.borrowQuery(); try { @@ -1687,7 +1794,8 @@ A query ends early in four ways: `QwpEgressQueryTimeoutError` and the client sends a cancel to QuestDB. - **Cancel.** `await query.cancel()` asks QuestDB to stop. Iteration and `completion` reject with `QwpEgressQueryError` whose `status` is `0x0a` - (CANCELLED). + (CANCELLED). A query that has already finished, or a DDL or DML statement, + completes normally instead. - **Leaving the loop.** `break`, `return`, or an exception inside `for await` cancels the query, and `completion` rejects with `QwpEgressQueryAbandonedError`. @@ -1695,6 +1803,21 @@ A query ends early in four ways: resolves `false` when the wait times out and leaves the query running. `query.isDone()` reports whether the query has ended. +QuestDB acts on a cancel between result batches, but while a query streams +without a [credit window](#flow-control), it may not read the cancel until the +whole result is sent. Without a credit window: + +- `cancel()` can end with `QwpEgressQueryCancelTimeoutError` instead of the + `0x0a` rejection: the client waits `query_close_timeout_ms` (5 seconds) for + QuestDB to stop, then closes the connection. +- After a deadline or an early exit from the loop, returning the lease can take + up to twice `query_close_timeout_ms`. + +Set a credit window when you cancel queries or use deadlines, with +`initial_credit` in the connect string or `initialCredit` per query. About +1 MiB is enough: the client replenishes it as your loop consumes batches, and +QuestDB then stops within milliseconds. + ```typescript import { connectQwpNodeClient, @@ -1707,7 +1830,8 @@ try { try { const query = await lease.query( "SELECT symbol, avg(price) FROM trades SAMPLE BY 1m", - { timeoutMs: 5_000 }, + // The credit window lets QuestDB act on the cancel promptly. + { timeoutMs: 5_000, initialCredit: 1024 * 1024 }, ); for await (const batch of query) { console.log(batch.rowCount); @@ -1727,9 +1851,10 @@ try { After a query ends early, its connection stays busy until QuestDB confirms the cancellation, and another `query()` on the same lease throws `a QWP query is already active on this connection`. Return the lease with -`close()` and borrow a new one for the next query. `close()` waits for the -cancellation, up to `query_close_timeout_ms` (5 seconds), and discards the -connection if QuestDB does not confirm in time. +`close()` and borrow a new one for the next query. `close()` waits up to +`query_close_timeout_ms` (5 seconds) for QuestDB to confirm the cancellation. +If it does not, `close()` discards the connection, which can take as long +again. ### Flow control @@ -1762,7 +1887,9 @@ try { ``` With `autoCredit: false`, call `query.grantCredit(bytes)` yourself. To cap the -rows in each batch, set `max_batch_rows` (1 to 1,048,576). +rows in each batch, set `max_batch_rows` (1 to 1,048,576). A credit window +also lets QuestDB act on a cancel or a deadline promptly; see +[Cancellation and timeouts](#cancellation-and-timeouts). ### Zero-copy result views @@ -1781,6 +1908,8 @@ try { const query = await lease.queryViews( "SELECT timestamp, symbol, price, amount FROM trades", (batch) => { + // A failover replays from batch 0; discard the failed attempt's sum. + if (batch.batchSequence === 0n) notional = 0; const price = batch.column(2); const amount = batch.column(3); for (let row = 0; row < batch.rowCount; row++) { @@ -1881,8 +2010,8 @@ try { .doubleColumn("price", 2615.54) .doubleColumn("amount", 0.5) .at(Date.now(), "ms"); - const sequence = await sender.flushAndGetSequence(); - await sender.waitForAcknowledged(sequence, 10_000); + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } finally { // Rethrows a terminal error. The pool then replaces the sender. await sender.close(); @@ -1907,7 +2036,7 @@ A standalone sender's `close()` can also reject, with | `appliedPolicy` | `string` | What the client did: `retriable` (reconnect and resend), `retriable-other` (resend to another endpoint), `terminal` (the sender stopped), or `abandoned` (journaled data was quarantined). | | `serverStatusByte` | `number` | The raw QWP status code, for example `0x03` for a schema mismatch. Absent for client-side errors. | | `serverMessage` | `string` | QuestDB's error text, for example `cannot parse DOUBLE from string [value=abc, column=price]`. | -| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `flushAndGetSequence()`. | +| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `publishedSequence`. | | `messageSequence` | `bigint` | The wire sequence of the rejected message. | | `tableName` | `string` | The table, when the server attributes the rejection to one. Often absent. | | `detectedAtMs` | `number` | When the client received the rejection. | @@ -1999,12 +2128,15 @@ connection). The lease remains usable after a `QwpEgressQueryError`. | Status | Name | Meaning | |---|---|---| -| `0x03` | SCHEMA_MISMATCH | A bind type is incompatible with its placeholder | -| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column | -| `0x06` | INTERNAL_ERROR | Server-side execution failure | +| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column, or a bind value the statement cannot use, such as a boolean for `LIMIT` | +| `0x06` | INTERNAL_ERROR | Execution failure, including a bind value that cannot be converted, such as `'abc'` compared with a DOUBLE column, and a column type that QWP cannot return | | `0x08` | SECURITY_ERROR | Missing permission | | `0x0a` | CANCELLED | The query was cancelled with `cancel()` | -| `0x0b` | LIMIT_EXCEEDED | A protocol limit was exceeded | +| `0x0b` | LIMIT_EXCEEDED | A server limit was reached: the server-side query timeout, memory, or a result row too large to send | + +The `QWP_STATUS` export names these codes, for example +`QWP_STATUS.PARSE_ERROR`, so code can compare against constants instead of +numbers. Other query errors: @@ -2032,6 +2164,43 @@ and has no server-side correlation ID beyond `requestId`. | `QwpDurableAckUnavailableError` | `request_durable_ack=on`, but the server does not support it. | | `QwpClientClosedError` | The pooled client, or a returned lease, is already closed. | +The pooled client wraps every failure to open a connection, from +`connectQwpNodeClient()`, `db.connect()`, `borrowSender()`, or +`borrowQuery()`, in a `QwpPoolResourceError`. Unwrap its `cause` before +checking for a specific error. When `addr` lists several hosts, the cause is a +`QwpFailoverError` whose `attempts` hold the error of each endpoint: + +```typescript +import { + connectQwpNodeClient, + QwpFailoverError, + QwpPoolResourceError, + QwpUpgradeError, +} from "@questdb/nodejs-client"; + +// The errors behind a failed connection, one per endpoint tried. +function connectionErrors(error: unknown): unknown[] { + const cause = error instanceof QwpPoolResourceError ? error.cause : error; + return cause instanceof QwpFailoverError + ? cause.attempts.map((attempt) => attempt.error) + : [cause]; +} + +try { + const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); + await db.close(); +} catch (error) { + for (const cause of connectionErrors(error)) { + if (cause instanceof QwpUpgradeError && cause.kind === "authentication") { + console.error("QuestDB rejected the credentials:", cause.message); + } else { + console.error("cannot connect:", cause); + } + } + throw error; +} +``` + An authentication failure (HTTP 401 or 403) ends the connection attempt for the whole endpoint list, because a credential rejected by one node is wrong for all of them, and it is not retried. The exception is a store-and-forward sender @@ -2068,8 +2237,10 @@ walks the list until it finds the current primary. Queries can use any node. never fall back to the primary, and they fail when no replica is reachable, including against a single open source server. Because the pooled client opens a query connection at startup, `connectQwpNodeClient()` then rejects too, with -`QwpPoolResourceError` caused by `QwpRoleMismatchError`. Set `query_pool_min=0` -to start without a replica. `zone` prefers endpoints in the same zone. +a `QwpPoolResourceError` whose `cause` is a `QwpRoleMismatchError`, or a +`QwpFailoverError` holding one per endpoint when `addr` lists several hosts. +Set `query_pool_min=0` to start without a replica. `zone` prefers endpoints in +the same zone. :::caution `target` in the connect string also filters ingestion @@ -2126,7 +2297,10 @@ endpoint when there is one, and runs the query again from the start: When the budget runs out, the query rejects with `QwpReconnectExhaustedError`. A `QwpEgressQueryError` from the server is a query result and never triggers -failover. +failover. Replaying an in-flight `query()` also re-executes DDL and DML: an +`INSERT` may run twice if its completion was lost. For non-idempotent SQL, +use a separate client configured with `failover=off` and check an uncertain +outcome before retrying; see [DDL and DML statements](#ddl-and-dml-statements). :::warning Clear partial results when a query restarts @@ -2209,7 +2383,13 @@ function onEvent(event: QwpReconnectEvent) { const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { // Each object replaces the reconnect_* or failover* keys from the connect // string. Fields left out use the built-in defaults. - ingressSession: { reconnect: { onEvent } }, + ingressSession: { + reconnect: { onEvent }, + // No event marks a terminal failure: it arrives here instead. + onError: (event) => { + if (event.terminal) console.error("ingestion stopped:", event.error); + }, + }, egressSession: { reconnect: { onEvent } }, }); await db.close(); @@ -2223,17 +2403,23 @@ failing on the first error. |---|---| | `connected` | The first connection succeeded. | | `reconnecting` | The active connection was lost. `cause` holds the error. | -| `attempt-failed` | One connection attempt failed. The client keeps trying. | +| `attempt-failed` | One connection attempt failed. The client retries if the error is retryable and its budget allows; otherwise this is the last event before the failure is reported. | | `reconnected` | Reconnected to the same endpoint. | | `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | -| `durable-ack-unavailable` | A store-and-forward sender is waiting for an endpoint that supports durable acknowledgement. | +| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. Only senders that retry indefinitely wait: store-and-forward senders after their first connection, and senders with `initial_connect_retry=async` or `lazy_connect=on`. | | `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | -| `primary-unavailable` | No reachable endpoint can currently accept writes. | +| `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | `reconnected` and `failed-over` are mutually exclusive: code that tracks the -current node must handle both. Neither `attempt-failed` nor -`primary-unavailable` is terminal: the client keeps retrying until its budget -runs out. +current node must handle both. + +No event marks a terminal failure. When a sender stops retrying, because its +reconnect budget ran out or the error cannot be retried, +`ingressSession.onError` receives the error with `terminal: true`, even while +the sender is idle. The sender's next `flush()`, auto-flushing `at()`, or +`close()` then rejects with the same error, such as +`QwpReconnectExhaustedError`. A query that cannot fail over rejects its +iteration and `completion` instead. For ingestion, `ingressSession` also accepts `onProgress`, for published, acknowledged, and durably acknowledged sequences, and `onError`, for session @@ -2417,13 +2603,19 @@ function logConnection(event: QwpReconnectEvent) { const db = await connectQwpNodeClient( "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + `token=${token};` + - "sender_pool_max=4;query_pool_max=8;", + // query_pool_min=0: start, and ingest, even while no replica is reachable. + "sender_pool_max=4;query_pool_min=0;query_pool_max=8;", { - // Queries run on replicas only; ingestion always follows the primary. + // Queries run on replicas only, never on the primary; ingestion always + // follows the primary. egress: { target: "replica", compression: "zstd" }, ingressSession: { onSenderError: (error: QwpSenderError) => console.error("batch rejected:", error.category, error.serverMessage), + // Terminal failures, such as an exhausted reconnect budget. + onError: (event) => { + if (event.terminal) console.error("ingestion stopped:", event.error); + }, // Replaces any reconnect_* keys; omitted fields use the defaults. reconnect: { onEvent: logConnection }, }, @@ -2453,8 +2645,8 @@ try { .doubleColumn("amount", amount) .at(Date.now(), "ms"); } - const sequence = await sender.flushAndGetSequence(); - await sender.waitForAcknowledged(sequence, 10_000); + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } catch (error) { if (!(error instanceof QwpIngressAckTimeoutError)) throw error; console.warn("rows not acknowledged yet; they stay queued for replay"); @@ -2462,7 +2654,8 @@ try { await sender.close(); } - // Querying: rows may not be visible yet, see "Read-after-write". + // Querying: rows may not be visible yet, see "Read-after-write". While no + // replica is reachable, borrowQuery() rejects with QwpPoolResourceError. const lease = await db.borrowQuery(); try { const query = await lease.query( diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index 05dd9ae724..2eec7f73fc 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -76,7 +76,9 @@ Why each line matters: It bounds only the **blocking** initial connect (`initial_connect_retry=on` / `sync`). Once a sender is running, the reconnect loop never consults it and retries a transport outage forever. Setting a large value here does nothing for -a running producer. See [Reconnect and outage handling](#reconnect-and-outage-handling). +a running producer. The Node.js client is the exception: a sender with neither +`sf_dir` nor `initial_connect_retry=async` applies it to every outage. See +[Reconnect and outage handling](#reconnect-and-outage-handling). ::: @@ -183,10 +185,10 @@ callers block up to `acquire_timeout_ms` then throw. | `sender_id` | `default` | | `sf_max_segment_bytes` (segment size) | `4 MiB` | | `sf_max_total_bytes` | `10 GiB` (SF mode) · `128 MiB` (memory mode) | -| `sf_durability` | `memory` (also supports `periodic`) | +| `sf_durability` | `memory` (also supports `periodic`; Node.js also `append`) | | `sf_sync_interval_millis` | `5000` (requires `sf_durability=periodic`) | | `sf_append_deadline_millis` | `30000` | -| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** | +| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** (Node.js memory mode: every outage) | | `reconnect_initial_backoff_millis` | `100` | | `reconnect_max_backoff_millis` | `5000` | | `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/Node.js) | @@ -210,7 +212,7 @@ callers block up to `acquire_timeout_ms` then throw. There is no "retry forever" setting to look for on the reconnect keys — a running sender already does. `reconnect_max_duration_millis` applies only to a -blocking initial connect; see +blocking initial connect, except on a Node.js sender in memory mode; see [Reconnect and outage handling](#reconnect-and-outage-handling). --- diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index 26edbbaa2c..e651e7a849 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1147,7 +1147,7 @@ section of the connect string reference: | Key | Default | Description | |----------------------------------|-----------|-------------------------------------------| -| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely. | +| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in memory mode. | | `reconnect_initial_backoff_millis` | `100` | First post-failure sleep. | | `reconnect_max_backoff_millis` | `5000` | Cap on per-attempt sleep. | | `initial_connect_retry` | `off` | Retry on first connect (`on`, `sync`, `async`). | @@ -1158,6 +1158,9 @@ Key behaviors: every host's zone tier is equivalent and selection is based on health state only. The `zone=` connect-string key is accepted but silently ignored, so a connect string shared with egress clients works unchanged on ingress. + The Node.js client is the exception: it applies `zone=` and `target=` to + ingress too, so `target=replica` in a shared connect string stops its + ingestion. - **Authentication errors are terminal** at any host (`401`/`403`). The reconnect loop does not continue past them. - **`421 + X-QuestDB-Role`** is a role reject: transient if the role is diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index be95b64a9b..a5519df0ef 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -80,7 +80,8 @@ the host re-advertises a different zone. `target=primary` collapses every host's zone tier to `Same` — writers must follow the primary regardless of geography. Ingress is currently zone-blind in both storage modes, so the `zone=` key is silently accepted on ingress -connections and only takes effect on egress. +connections and only takes effect on egress. The Node.js client is the +exception: it applies `zone=` and `target=` to ingress too. ### Selection priority @@ -146,6 +147,9 @@ buffer capacity (`sf_max_total_bytes` and disk), not a timer. - Maximum backoff: `5 s` - Per-outage budget: **none**. `reconnect_max_duration_millis` bounds only the blocking sync initial connect, and the running loop never consults it. + The Node.js client is the exception: in memory mode, a sender without + `initial_connect_retry=async` gives up after `reconnect_max_duration_millis` + and fails with `QwpReconnectExhaustedError`. - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps reconnect storms when many producers share a cluster - Inter-host pause within a round: **none** — the client walks the full @@ -197,7 +201,8 @@ within the same round. No exponential backoff is consumed. - `421` + `X-QuestDB-Role: PRIMARY_CATCHUP` → `TransientReject` - `421` + any other non-empty role, including unrecognised tokens → `TopologyReject` -- `SERVER_INFO.Role` does not match the requested `target=` (egress only) +- `SERVER_INFO.Role` does not match the requested `target=` (egress only; the + Node.js client also applies it to ingress) If every host in a round role-rejects, ingress pays one fixed backoff sleep (reset to `InitialBackoff`, no doubling) and starts a fresh round; egress diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index ab669475f7..04ffd35ba2 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -17,6 +17,7 @@ first. `addr` and `auth_timeout_ms` apply to every WS / WSS / HTTP / HTTPS client. `zone` is accepted everywhere but only takes effect on egress; `target` is an egress-only key and is rejected as an unknown key on an ingress connect string. +The Node.js client is the exception: it applies both keys to ingress too. They are documented in full on the [connect-string reference](/docs/connect/clients/connect-string#failover-keys); the table below summarises the failover-relevant subset. @@ -47,7 +48,7 @@ for the full list. The failover-relevant keys are: | Key | Type | Default | Notes | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender with neither `sf_dir` nor `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | @@ -61,7 +62,7 @@ network), and retrying for five minutes only hides it. | Value | Behaviour | |---|---| | `off` (default; alias `false`) | First-connect failure is terminal. The producer's call to build the sender throws immediately. | -| `on` (aliases `sync`, `true`) | First-connect failures are retried on the caller's thread. The constructor blocks until it connects or `reconnect_max_duration_millis` expires — this is the **only** place that key applies. Once the sender is running, reconnection is unbounded. | +| `on` (aliases `sync`, `true`) | First-connect failures are retried on the caller's thread. The constructor blocks until it connects or `reconnect_max_duration_millis` expires — this is the **only** place that key applies. Once the sender is running, reconnection is unbounded. A Node.js sender in memory mode is the exception; see the `reconnect_max_duration_millis` row above. | | `async` | The constructor returns immediately; the background I/O thread drives the reconnect loop. The producer experiences backpressure if it tries to publish before the connection comes up. Intended for unattended producers where the SF directory may already carry segments from a prior process and the server may come up later. | ## Egress (query) diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 948aebf552..77a1709420 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -38,7 +38,10 @@ SF runs in either of two modes selected by the connect string: Both modes share the same reconnect loop, the same backoff and retry budgets, and the same on-the-wire behaviour. The only difference is -where unacked data lives. +where unacked data lives. The Node.js client is the exception: in memory +mode, a sender without `initial_connect_retry=async` gives up after +`reconnect_max_duration_millis` (5 minutes by default); see the +[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). ## What "frame" means here @@ -143,7 +146,9 @@ When the wire connection breaks — for any reason — the I/O thread enters the reconnect loop documented in [Client failover concepts](/docs/high-availability/client-failover/concepts/). The producer is **not notified**: it keeps publishing into the substrate, -bounded by `sf_max_total_bytes` (see backpressure below). +bounded by `sf_max_total_bytes` (see backpressure below). On the Node.js +client in memory mode, `flush()` instead waits for the reconnect, up to +`reconnect_max_duration_millis`. On every successful (re)connect: @@ -191,6 +196,10 @@ fires, a `WARN` is logged and: next sender on the same slot; - in **memory mode**, the un-acked tail is lost. +On the Node.js client, a standalone sender's `close()` rejects with +`QwpSenderCloseTimeoutError` instead of logging, and the pooled client's +`db.close()` reports the timeout to its `onError` callback. + Setting `close_flush_timeout_millis=0` (or `-1`) skips the drain wait entirely — useful for fast shutdown paths where you do not want to block. Even in this branch, the slot lock is released and segments are unmapped diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index a3712a102c..3ea50b925d 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -28,7 +28,7 @@ mode. | `sender_id` | string | `default` | Slot subdirectory name. Two senders sharing the same `sender_id` and `sf_dir` will collide on the slot lock. Must not contain path separators or be empty. | | `sf_max_segment_bytes` | size | `4M` | Per-segment file size; rotation threshold. | | `sf_max_total_bytes` | size | `128M` (memory) / `10G` (SF) | Hard cap on resident SF storage. Triggers producer backpressure when full. | -| `sf_durability` | enum | `memory` | `memory` (page-cache durable) and `periodic` (background checkpoint to stable storage) both ship. `periodic` requires `sf_dir`. `flush` and `append` parse but are rejected at build time. The .NET client accepts `memory` only. | +| `sf_durability` | enum | `memory` | `memory` relies on the page cache; `periodic` checkpoints in the background and requires `sf_dir`. Node.js also supports `append`, which makes each journal append durable before `flush()` resolves. Go and .NET accept only `memory`; other clients reject `append` at build time. `flush` is not supported. | | `sf_sync_interval_millis` | int (ms) | `5000` | Checkpoint cadence for `sf_durability=periodic`; rejected without it. A floor, not a guarantee: scheduler and storage latency add to it. | | `sf_append_deadline_millis` | int (ms) | `30000` | How long a producer `appendBlocking` call waits for ACK-driven trim to free space before throwing. | | `drain_orphans` | bool | `off` | Scan `/*` at startup and spawn drainers for sibling slots that contain unacked data. See [orphan adoption](/docs/high-availability/store-and-forward/concepts/#orphan-adoption). | @@ -48,7 +48,7 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender without `sf_dir` or `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender with neither `sf_dir` nor `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | @@ -90,12 +90,12 @@ canonical entries. | `username` / `password` | string | unset | HTTP Basic auth on the upgrade request. | | `token` | string | unset | Bearer token on the upgrade request. | | `tls_verify` | enum | `on` | `on` or `unsafe_off`. Applies to `wss::` / TLS connections. | -| `tls_roots` | path | system trust | Custom CA trust store. | +| `tls_roots` | path | system trust (Node.js: bundled CAs) | Custom CA trust store. | | `tls_roots_password` | string | unset | Trust store password. | | `auto_flush` | bool | `on` | Global on/off for auto-flush triggers. | | `auto_flush_rows` | int / `off` | `1000` | Row-count flush trigger. | | `auto_flush_bytes` | int / `off` | `0` (off) | Byte-size flush trigger. | -| `auto_flush_interval` | int (ms) / `off` | `100` | Time-since-first-row flush trigger. | +| `auto_flush_interval` | int (ms) / `off` | `100` | Time-since-first-row flush trigger (Node.js: since last flush or sender creation). | | `init_buf_size` | size | `64K` | Initial encode buffer capacity. | | `max_buf_size` | size | `100M` | Max encode buffer capacity. | | `max_name_len` | int | `127` | Local validation cap for table / column names. | @@ -106,8 +106,9 @@ The parser rejects: - Unknown keys (forward compatibility is via the spec, not silent acceptance). -- `sf_durability` values other than `memory`, `flush`, `append`. `flush` - and `append` parse but are rejected at build time today. +- Unsupported `sf_durability` values. Go and .NET accept only `memory`; + Node.js also accepts `periodic` and `append`. Other clients accept + `periodic`, but reject `flush` and `append` at build time. - `sender_id` containing path separators or empty. - `request_durable_ack=on` on non-WebSocket transports. diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index 9b69e434ca..d6a24df0a2 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -20,8 +20,9 @@ In SF mode every sender owns one **slot directory**: ``` // -├── .lock # advisory exclusive lock (kernel-released on process exit) +├── .lock # OS lock; Node.js retains it for compatibility only ├── .lock.pid # UTF-8 text: holder PID + '\n' (diagnostic only) +├── .lock.owner/ # Node.js ownership directory; may survive a crash ├── .failed # optional drainer-failure sentinel (UTF-8 reason text) ├── .ack-watermark # optional 16-byte durable-ack high-water mark ├── sf-0000000000000001.sfa @@ -36,10 +37,10 @@ the host. ### `.lock` and `.lock.pid` -The `.lock` file is held under an advisory exclusive lock for the engine's -lifetime — POSIX clients use `flock` / `fcntl`, Windows uses -`LockFileEx`. The lock is released automatically when the file descriptor -closes, including on hard process exit (kernel cleanup). +For clients other than Node.js, the `.lock` file is held under an advisory +exclusive lock for the engine's lifetime — POSIX clients use `flock` / +`fcntl`, Windows uses `LockFileEx`. Their lock is released automatically when +the file descriptor closes, including on hard process exit (kernel cleanup). A second sender pointing at the same slot directory will fail to start with an error that names the holder's PID, read from `.lock.pid`. The @@ -53,6 +54,22 @@ files are harmless — the next acquirer silently overwrites them. **not** share a slot on a network filesystem. Their lock primitives are incompatible. +:::caution Node.js client + +The Node.js client does not use an OS lock. It locks a slot by creating a +`.lock.owner` directory inside it, which records the owner's host name and +process ID, and keeps `.lock` and `.lock.pid` only for compatibility. A +crashed Node.js sender therefore leaves the slot locked: a new sender takes it +over automatically only on the same host, once the recorded process ID is no +longer in use. That often fails in containers, where the application usually +runs as process ID 1 and a replacement container has a new host name. +Otherwise, make sure the previous process has exited, then delete +`.lock.owner`. Node.js and other clients do not see each other's locks, so +never let them use the same `sf_dir` at the same time. See the +[Node.js client](/docs/connect/clients/nodejs/#store-and-forward). + +::: + ### `.failed` Present iff a previous drainer attempt gave up on the slot — reconnect @@ -89,8 +106,10 @@ and the second start fails loudly. A common cause is a redeploy where the old process hasn't fully exited when the new one comes up. Solutions: -- Wait for the old process to release the lock (the kernel releases on - exit; `kill -9` is sufficient). +- Stop the previous process. For clients using OS locks, the kernel releases + the lock on exit (even after `kill -9`). A killed Node.js sender can leave + `.lock.owner` behind: verify the old owner is gone before removing it; see + [`.lock` and `.lock.pid`](#lock-and-lockpid). - Use a deployment unit that orders shutdown before startup. - For containerised deployments, set `sender_id` from a per-pod stable identity so two pods with the same template name don't collide. @@ -182,7 +201,7 @@ fresh start: no segments, no replay. | Symptom | Likely cause | Operator action | |---|---|---| -| "Slot held by PID ``" | Two processes claiming the same `sender_id`. | Stop the duplicate. The lock releases on its exit. | +| "Slot held by PID ``" or `QwpReplayStoreLockedError` (Node.js) | Another process holds the slot, or a Node.js `.lock.owner` is stale after a crash. | Stop the duplicate. OS locks release on exit; for Node.js verify the owner is gone before removing `.lock.owner` (see [`.lock` and `.lock.pid`](#lock-and-lockpid)). | | "Gap between segments" | Corruption — a segment was deleted out of band. | Restore from backup or accept data loss; the substrate refuses to start. | | "Watermark exceeds publishedFsn" | `.ack-watermark` is corrupt; the engine falls back to the no-watermark seed. | Logged as `WARN`. Replay will re-send the lowest segment's frames; rely on server deduplication. | | Torn tail count > 0 | The previous process crashed mid-frame-write. | Informational; the CRC + zero-fill design discards the partial frame. | @@ -197,6 +216,9 @@ fresh start: no segments, no replay. | `0` or `-1` | Skip the drain wait. Pending data persists on disk (SF) for the next sender, or is lost (memory). | | any other positive value | That timeout in milliseconds. | +On the Node.js client, a standalone sender's `close()` rejects with +`QwpSenderCloseTimeoutError` on timeout instead of logging a `WARN`. + In every branch `close()`: - Performs a non-blocking safety-net check that rethrows any latched diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 8897ed6a9e..27cbc02827 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -57,7 +57,9 @@ Unacked frames are written to mmap'd files under Both modes share the same wire behaviour, the same failover loop, and the same connect-string keys for everything other than storage. You can switch between them without changing application code — only the connect -string. +string. On the Node.js client, memory mode also gives up after +`reconnect_max_duration_millis` of outage, unless `initial_connect_retry=async` +is set. ## Comparison at a glance diff --git a/documentation/query/overview.md b/documentation/query/overview.md index e39f890049..469e3db7c2 100644 --- a/documentation/query/overview.md +++ b/documentation/query/overview.md @@ -97,19 +97,24 @@ against the demo instance. ## QuestDB client libraries The official client libraries speak the QuestDB Wire Protocol (QWP), a binary -protocol that carries both ingestion and query traffic over one connection and -one configuration string. This is the fastest way to get data out of QuestDB -from an application. +protocol for both ingestion and querying, configured with one connection +string. A client may use separate connections for those operations: the +[Node.js client](/docs/connect/clients/nodejs/#the-connection-pool), for example, +maintains separate sender and query pools. Results stream rather than arriving in one block. The server sends batches as it produces them, so an application starts processing the head of a result while the tail is still being computed, and a result larger than memory never has to be materialized at all. -Connections recover on their own. When a connection drops and replicas are -available, the client reconnects and retries against another one without the -application intervening. A query that fails over restarts from the beginning, -which is transparent if you materialize the whole result. +Connections can recover from transport failures. With failover enabled, a +client may reconnect and re-execute an in-flight query, including on the same +host. Result rows then restart from the beginning. Some clients' materializers +discard their partial result automatically; if you process batches yourself, +reset any accumulated state on replay or handle the client's terminal error. +See [Node.js query failover](/docs/connect/clients/nodejs/#query-failover) for +an example. Re-execution can also repeat SQL writes; see +[DDL and DML statements](/docs/connect/clients/nodejs/#ddl-and-dml-statements). The Rust, C++, and Python clients hand back results as Arrow record batches. That is the native memory layout of diff --git a/plugins/raw-markdown/convert-components.js b/plugins/raw-markdown/convert-components.js index fc576b441f..29723cc450 100644 --- a/plugins/raw-markdown/convert-components.js +++ b/plugins/raw-markdown/convert-components.js @@ -605,18 +605,45 @@ function bumpHeadings(markdown, bumpBy = 1) { } /** - * Removes import statements from processed markdown - * Handles both single-line and multi-line imports + * Removes MDX import statements without touching imports in fenced examples. + * Handles both single-line and multi-line imports outside code fences. * @param {string} content - The markdown content - * @returns {string} Content with imports removed + * @returns {string} Content with MDX imports removed */ function removeImports(content) { - let processed = content - // First handle single-line imports - processed = processed.replace(/^import\s+.+\s+from\s+['"].+['"];?\s*$/gm, '') - // Then handle multi-line imports (where line breaks exist) - processed = processed.replace(/^import\s+[\s\S]*?\s+from\s*\n?\s*['"].+['"];?\s*$/gm, '') - return processed + const lines = content.split('\n') + const output = [] + let segmentStart = 0 + let fenceChar = '' + let fenceLen = 0 + + function appendOutside(end) { + if (segmentStart === end) return + let segment = lines.slice(segmentStart, end).join('\n') + segment = segment.replace(/^import\s+.+\s+from\s+['"].+['"];?\s*$/gm, '') + segment = segment.replace(/^import\s+[\s\S]*?\s+from\s*\n?\s*['"].+['"];?\s*$/gm, '') + output.push(...segment.split('\n')) + } + + for (let i = 0; i < lines.length; i++) { + const line = lines[i] + const fence = line.match(/^ {0,3}(`{3,}|~{3,})(.*)$/) + if (!fenceChar) { + if (!fence) continue + appendOutside(i) + fenceChar = fence[1][0] + fenceLen = fence[1].length + } else if (fence && fence[1][0] === fenceChar && + fence[1].length >= fenceLen && /^\s*$/.test(fence[2])) { + fenceChar = '' + fenceLen = 0 + segmentStart = i + 1 + } + output.push(line) + } + + if (!fenceChar) appendOutside(lines.length) + return output.join('\n') } /** diff --git a/plugins/raw-markdown/convert-components.test.js b/plugins/raw-markdown/convert-components.test.js new file mode 100644 index 0000000000..f3553e807a --- /dev/null +++ b/plugins/raw-markdown/convert-components.test.js @@ -0,0 +1,41 @@ +const assert = require('node:assert/strict') +const fs = require('node:fs') +const path = require('node:path') +const test = require('node:test') +const matter = require('gray-matter') +const { removeImports } = require('./convert-components') + +test('removes MDX imports but keeps TypeScript imports in fenced examples', () => { + const markdown = [ + 'import Widget from "@site/src/components/Widget"', + 'import {', + ' Tabs,', + ' TabItem,', + '} from "@theme/Tabs"', + '', + '```typescript', + 'import {', + ' connectQwpNodeClient,', + ' QwpEgressQueryError,', + '} from "@questdb/nodejs-client";', + '```', + '', + '~~~ts', + 'import { client } from "@questdb/nodejs-client";', + '~~~', + ].join('\n') + + const result = removeImports(markdown) + assert.doesNotMatch(result, /@site\/src\/components\/Widget|@theme\/Tabs/) + assert.match(result, /```typescript\nimport \{\n connectQwpNodeClient,\n QwpEgressQueryError,\n\} from "@questdb\/nodejs-client";\n```/) + assert.match(result, /~~~ts\nimport \{ client \} from "@questdb\/nodejs-client";\n~~~/) +}) + +test('preserves imports from the Node.js Quick start while removing its MDX import', () => { + const file = path.join(__dirname, '../../documentation/connect/clients/nodejs.md') + const { content } = matter(fs.readFileSync(file, 'utf8')) + const result = removeImports(content) + + assert.doesNotMatch(result, /^import SfDedupWarning from /m) + assert.match(result, /```typescript\nimport \{\n connectQwpNodeClient,\n QwpEgressQueryError,\n\} from "@questdb\/nodejs-client";/) +}) From 427bf4846e092ddb36deb1d968d9ee6ddffc5034 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 18:30:35 +0100 Subject: [PATCH 08/25] docs(nodejs): correct QWP examples and shared guidance --- .../connect/clients/connect-string.md | 8 +- documentation/connect/clients/nodejs.md | 185 ++++++++++++------ .../connect/compatibility/pgwire/nodejs.md | 11 +- documentation/connect/overview.md | 3 +- .../client-failover/configuration.md | 2 +- .../store-and-forward/concepts.md | 45 +++-- .../store-and-forward/configuration.md | 11 +- 7 files changed, 176 insertions(+), 89 deletions(-) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index be80fa99ac..f084e59686 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -364,7 +364,9 @@ applications must call `flush()` explicitly; see the *Applies to: ingress (encode buffer).* These keys control the in-memory row buffer that the client uses before -flushing. +flushing. The Node.js QWP `ws`/`wss` client rejects `init_buf_size` and +`max_buf_size` as legacy-transport options; use +[`auto_flush_rows`](#auto-flush) to control its row batches instead. - `init_buf_size` — initial buffer size in bytes. Default: `65536` (64 KiB). Accepts [size suffixes](#size-suffixes). @@ -909,7 +911,7 @@ description and behaviour notes. | `failover_backoff_max_ms` | int (ms) | `1000` | [Egress failover](#reconnect-keys) | | `failover_max_attempts` | int | `8` | [Egress failover](#reconnect-keys) | | `failover_max_duration_ms` | int (ms) | `30000` | [Egress failover](#reconnect-keys) | -| `init_buf_size` | size | `65536` (64 KiB) | [Buffer sizing](#buffer) | +| `init_buf_size` | size | `65536` (Node.js QWP: unsupported) | [Buffer sizing](#buffer) | | `initial_connect_retry` | enum (`off` / `on` / `async`) | `off` (auto-promoted to `on` when any explicit `reconnect_*` key is set) | [Ingress reconnect](#reconnect-keys) | | `initial_credit` | int (bytes) | `0` (unbounded) | [Query client keys](#egress-keys) | | `housekeeper_interval_ms` | int (ms) | `5000` | [Connection pool](#pool-keys) | @@ -918,7 +920,7 @@ description and behaviour notes. | `max_background_drainers` | int | `4` | [Store-and-forward](#sf-keys) | | `max_batch_rows` | int (`1`–`1048576`) | server default | [Query client keys](#egress-keys) | | `max_lifetime_ms` | int (ms) | `1800000` (`0` ⇒ infinite) | [Connection pool](#pool-keys) | -| `max_buf_size` | size | `104857600` (100 MiB) | [Buffer sizing](#buffer) | +| `max_buf_size` | size | `104857600` (Node.js QWP: unsupported) | [Buffer sizing](#buffer) | | `max_datagram_size` | size | (UDP) below typical MTU | [Buffer sizing](#buffer) | | `max_name_len` | int | `127` | [Buffer sizing](#buffer) | | `max_frame_rejections` | int (≥ 1) | `4` | [Error handling](#error-handling) | diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 15588195b7..9ac957b6ab 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -61,11 +61,14 @@ complete API from the package root, ships ES module and CommonJS builds, and bundles TypeScript declarations. There are no other supported import paths. The examples on this page are TypeScript ES modules with top-level `await`. -They also run as plain JavaScript once type annotations are removed. +To run them as plain JavaScript, use ES modules (`.mjs` or `"type": "module"`) +and remove type annotations, type-only imports, and TypeScript assertions +such as `as const`. ## Quick start -Connect with one connect string, write two rows, and read them back: +Connect with one connect string, write two rows, and try to query the ETH-USD +row. Ingestion is asynchronous, so an immediate read may not see it yet: ```typescript import { @@ -152,57 +155,74 @@ microseconds since the Unix epoch; see When `flush()` resolves, the client has published the rows, but QuestDB may not have received them yet. QuestDB acknowledges a batch once it has committed it to its write-ahead log, and applies committed rows to the table -asynchronously. A query that runs right after -ingestion can therefore fail with `table does not exist` on a first run, or -succeed and return no rows. When your code must read its own writes, poll until -the rows appear, bounded by a deadline: +asynchronously. A query that runs right after ingestion can therefore fail +with `table does not exist` on a first run, or succeed and return no rows. + +When your code must read its own writes, create the table first, write an event +with a unique ID, and poll for **that ID**. Pre-creating the table avoids +mistaking an unrelated SQL error for the first-write table-creation delay. Give +each query the time remaining until the deadline so a stalled query cannot +leave the poll running indefinitely: ```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, - type QwpClient, -} from "@questdb/nodejs-client"; +import { randomUUID } from "node:crypto"; +import { connectQwpNodeClient } from "@questdb/nodejs-client"; -async function countRows(db: QwpClient, sql: string): Promise { +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { const lease = await db.borrowQuery(); try { - const query = await lease.query(sql); - let rows = 0; - for await (const batch of query) rows += batch.rowCount; - await query.completion; - return rows; - } finally { - await lease.close(); - } -} + const ddl = await lease.query( + "CREATE TABLE IF NOT EXISTS trades_readback (" + + "timestamp TIMESTAMP, trade_id VARCHAR, symbol SYMBOL" + + ") TIMESTAMP(timestamp) PARTITION BY DAY", + ); + await ddl.completion; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sql = "SELECT * FROM trades WHERE symbol = 'ETH-USD' LIMIT 10"; - const deadline = Date.now() + 10_000; - let rows = 0; - while (rows === 0) { + const tradeId = randomUUID(); + const sender = await db.borrowSender(); try { - rows = await countRows(db, sql); - } catch (error) { - // The table may not exist yet: keep polling until the deadline. - if (!(error instanceof QwpEgressQueryError) || Date.now() >= deadline) { - throw error; - } + await sender + .table("trades_readback") + .stringColumn("trade_id", tradeId) + .symbol("symbol", "ETH-USD") + .at(Date.now(), "ms"); + await sender.flush(); + } finally { + await sender.close(); } - if (rows === 0) { - if (Date.now() >= deadline) throw new Error("rows not visible in time"); - await new Promise((resolve) => setTimeout(resolve, 100)); + + const deadline = Date.now() + 10_000; + let visible = false; + while (!visible) { + const remainingMs = deadline - Date.now(); + if (remainingMs <= 0) throw new Error("trade not visible in time"); + const query = await lease.query( + "SELECT trade_id FROM trades_readback WHERE trade_id = $1 LIMIT 1", + { + binds: (binds) => binds.setVarchar(0, tradeId), + timeoutMs: remainingMs, + }, + ); + for await (const batch of query) visible ||= batch.rowCount > 0; + await query.completion; + if (!visible) { + await new Promise((resolve) => + setTimeout(resolve, Math.min(100, Math.max(0, deadline - Date.now()))), + ); + } } + console.log(`visible trade: ${tradeId}`); + } finally { + await lease.close(); } - console.log(`visible rows: ${rows}`); } finally { await db.close(); } ``` -Do not replace the poll with a fixed sleep: the apply latency varies with load. +SQL errors now surface instead of being retried. Do not replace the poll with a +fixed sleep: the apply latency varies with load. ## Connecting @@ -647,7 +667,10 @@ until the first query. import { connectQwpNodeClient } from "@questdb/nodejs-client"; // Resolves immediately, even if QuestDB is not running yet. -const db = await connectQwpNodeClient("ws::addr=localhost:9000;lazy_connect=on;"); +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;lazy_connect=on;" + + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=startup-a;sf_durability=append;", +); try { const sender = await db.borrowSender(); try { @@ -666,6 +689,10 @@ try { } ``` +Use a writable, persistent `sf_dir` and reuse the same `sender_id` after a +restart so the example's rows survive shutdown while QuestDB is down. Without +`sf_dir`, the example would discard them when `db.close()` finishes. + `lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, and rejects an explicit conflicting value. Setting `initial_connect_retry=async` without `lazy_connect` is not enough: the query pool still connects at startup, @@ -1229,8 +1256,11 @@ try { await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } catch (error) { if (error instanceof QwpIngressAckTimeoutError) { - // Not acknowledged in time. The rows are still pending, not lost. - console.warn("ACK timeout at", error.acknowledgedSequence); + // Still pending in memory, but closing without sf_dir can lose them. + console.warn( + "ACK timeout; rows may be lost on close at", + error.acknowledgedSequence, + ); } else { throw error; } @@ -2505,10 +2535,10 @@ connect string and calling `connect()`: | Auto-flush interval | 1,000 ms | 100 ms | | `flush()` completes when | QuestDB responds to the HTTP request | The batch is published; the ACK arrives later | | Server rejection | `flush()` throws | Asynchronous: `onSenderError`, `waitForAcknowledged()`, or `flush()` with `awaitServerAck` | -| Rows staged at `close()` | Lost unless flushed | Published, then acknowledged within 5 seconds | +| Rows staged at `close()` | Lost unless flushed | Published; waits up to 5 seconds for ACK, then unacknowledged rows may be lost without `sf_dir` | | Reconnect and replay | Retries one request for `retry_timeout` | Automatic, with replay of unacknowledged batches | | Store-and-forward, querying, pooling | Not available | Available | -| Column types | ILP types | Every QuestDB type | +| Column types | ILP types | More types, subject to [column-method](#column-methods) and [array](#arrays) support | Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, `max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss`, @@ -2578,8 +2608,26 @@ and the [ILP overview](/docs/connect/compatibility/ilp/overview/). ## Full example: ingestion and querying with failover -A production service that ingests trades and queries recent prices, with TLS, -a token, several hosts, error handling, and failover handling: +A production-oriented pattern that ingests trades and queries recent prices, +with TLS, a token, several hosts, error handling, and failover handling. Before +running it, create the deduplicated table on the primary (or reuse the table +from [Store-and-forward](#store-and-forward)): + +```questdb-sql +CREATE TABLE IF NOT EXISTS trades_sf ( + timestamp TIMESTAMP, + trade_id SYMBOL, + symbol SYMBOL, + side SYMBOL, + price DOUBLE, + amount DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, trade_id); +``` + +Replace the sample events with source-assigned trade IDs and timestamps. Keep +both values unchanged when retrying the same event, and use a writable, +persistent `sf_dir` so unacknowledged rows survive a shutdown: ```typescript import { @@ -2604,7 +2652,8 @@ const db = await connectQwpNodeClient( "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + `token=${token};` + // query_pool_min=0: start, and ingest, even while no replica is reachable. - "sender_pool_max=4;query_pool_min=0;query_pool_max=8;", + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + + "sf_durability=append;sender_pool_max=4;query_pool_min=0;query_pool_max=8;", { // Queries run on replicas only, never on the primary; ingestion always // follows the primary. @@ -2630,26 +2679,41 @@ const db = await connectQwpNodeClient( ); try { - // Ingestion: one borrowed sender per producer. + // Ingestion: one borrowed sender per producer. IDs and timestamps must + // come from the source, not be regenerated on an application retry. + const events = [ + { + tradeId: "trade-12345", + timestampMs: 1723000000000, + symbol: "ETH-USD", + price: 2615.54, + amount: 0.5, + }, + { + tradeId: "trade-12346", + timestampMs: 1723000000001, + symbol: "BTC-USD", + price: 39269.98, + amount: 0.001, + }, + ]; const sender = await db.borrowSender(); try { - for (const [symbol, price, amount] of [ - ["ETH-USD", 2615.54, 0.5], - ["BTC-USD", 39269.98, 0.001], - ] as const) { + for (const event of events) { await sender - .table("trades") - .symbol("symbol", symbol) + .table("trades_sf") + .symbol("trade_id", event.tradeId) + .symbol("symbol", event.symbol) .symbol("side", "buy") - .doubleColumn("price", price) - .doubleColumn("amount", amount) - .at(Date.now(), "ms"); + .doubleColumn("price", event.price) + .doubleColumn("amount", event.amount) + .at(event.timestampMs, "ms"); } await sender.flush(); await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } catch (error) { if (!(error instanceof QwpIngressAckTimeoutError)) throw error; - console.warn("rows not acknowledged yet; they stay queued for replay"); + console.warn("ACK timed out; rows remain in sf_dir for replay after close"); } finally { await sender.close(); } @@ -2659,7 +2723,7 @@ try { const lease = await db.borrowQuery(); try { const query = await lease.query( - "SELECT timestamp, symbol, price FROM trades " + + "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, ); @@ -2682,6 +2746,11 @@ try { } ``` +The query can still miss newly acknowledged rows until WAL apply catches up; +use the [Read-after-write](#read-after-write) pattern for a visibility guarantee. +A replayed batch is idempotent only because this example retains the event's +ID and timestamp and enables table-level deduplication. + ## Next steps - [Connect string reference](/docs/connect/clients/connect-string/) for every diff --git a/documentation/connect/compatibility/pgwire/nodejs.md b/documentation/connect/compatibility/pgwire/nodejs.md index bc9123385e..4d02825bcf 100644 --- a/documentation/connect/compatibility/pgwire/nodejs.md +++ b/documentation/connect/compatibility/pgwire/nodejs.md @@ -37,6 +37,15 @@ standard PostgreSQL driver or ORM. ::: +:::note Example schema + +The PGWire examples below assume a pre-existing `trades` table with `ts` as +its designated timestamp and `symbol` and `price` columns. This differs from +the [QWP Node.js quick start](/docs/connect/clients/nodejs/#quick-start), which +auto-creates `trades.timestamp`. To query that table with these examples, +replace SQL `ts` and JavaScript `.ts` with `timestamp` and `.timestamp`. + +::: ## Connection Parameters @@ -748,7 +757,7 @@ async function latestByQuery() { // Get the latest values for each symbol const latest = await sql` SELECT * FROM trades - LATEST ON timestamp PARTITION BY symbol + LATEST ON ts PARTITION BY symbol ` console.log(`Latest prices for ${latest.length} symbols:`) diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 941d3df250..6eda395eea 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -31,7 +31,8 @@ Pick the path that matches your environment. The first-party libraries for **Java, Python, Go, Rust, Node.js, C & C++, and .NET** are the recommended way to talk to QuestDB. They speak the **QuestDB Wire Protocol (QWP)** and unify ingest and query under one -configuration and one connection. +client configuration. The client may use separate connections for ingestion +and queries, as the Node.js library does. ### QWP support diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 04ffd35ba2..ffa360d4c3 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -27,7 +27,7 @@ the table below summarises the failover-relevant subset. | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | | `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | | `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. The [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | -| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover the TCP connect or TLS handshake — those use the OS default. Set lower if you have well-known network paths and want faster failover; set higher only if upgrade is genuinely slow. | +| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. The Node.js client bounds DNS and TCP/TLS separately with `connect_timeout` (15 s by default); other clients may use the OS default. Lower it for faster upgrade failure detection; tune the connect timeout separately. | `addr` syntax — both of these are equivalent and produce the same three-peer list: diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 77a1709420..192c0e5cc3 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -85,9 +85,13 @@ Two consequences: - Frames **must** be sent in strict order. The wire format does not serialise `wireSeq` — the server assigns it implicitly from receive order. Reordering breaks the FSN mapping. -- After a reconnect, the server sees the **same payloads** at new - `wireSeq` values. Server-side dedup keys off `messageSequence` inside - the payload, not `wireSeq`, so replay does not produce double-writes. +- After a reconnect, the server may see the **same payloads** at new + `wireSeq` values. An acknowledgement can be lost after a batch was + committed, so replay is **at least once** and can insert duplicate rows. + `wireSeq` is transport bookkeeping, not a deduplication key. For + idempotent ingestion, use stable source IDs and timestamps with table-level + `DEDUP UPSERT KEYS`, as in the [Node.js store-and-forward + example](/docs/connect/clients/nodejs/#store-and-forward). ## Trim: how unacked data is reclaimed @@ -308,25 +312,26 @@ shared `sf_dir`, blindly draining unknown slots may be surprising. ## Error frames -Not every server response is an OK. Server errors fall into six -categories, each with a default policy: +Not every server response is an OK. A rejected batch is **not** silently +dropped and trimmed: the client either retries it or reports a terminal error. +The defaults below apply to the Node.js QWP client; consult the +[connect-string error policies](/docs/connect/clients/connect-string/#error-handling) +for the shared vocabulary and per-client override support. -| Category | Default | Meaning | +| Category | Node.js default | Meaning | |---|---|---| -| `SCHEMA_MISMATCH` | `DROP_AND_CONTINUE` | The batch's schema doesn't match the server. Replay won't help — the substrate logs and advances trim past the rejected span. | -| `WRITE_ERROR` | `DROP_AND_CONTINUE` | Per-batch write failure (e.g. table is not currently accepting writes). | -| `PARSE_ERROR` | `HALT` | Almost certainly a client bug. The substrate preserves on-disk frames for postmortem. | -| `INTERNAL_ERROR` | `HALT` | Catch-all server fault. | -| `SECURITY_ERROR` | `HALT` | Cluster-wide auth / authorization failure. | -| `PROTOCOL_VIOLATION` | `HALT` (forced) | Connection is gone after a terminal WebSocket close code; no choice. | - -Errors are also delivered to an **error inbox** — a bounded queue -consumed by a daemon dispatcher that invokes your registered handler. -Overflow drops the oldest entry rather than the newest (watermarks are -monotonic; the latest entry is the most informative). The default -handler logs every received error: silence is forbidden by the contract, -because a buggy or no-op handler would hide data loss -indistinguishably from a healthy connection. +| `SCHEMA_MISMATCH` | `terminal` | The schema does not match. The sender stops; in SF mode, the rejected bytes remain in the journal for inspection. | +| `WRITE_ERROR` | `retriable` | A write failed (for example, a temporary storage problem); reconnect and replay. | +| `PARSE_ERROR` | `terminal` | Malformed payload; replaying identical bytes cannot help. | +| `INTERNAL_ERROR` | `retriable` | Retry after an unexpected server-side failure. | +| `SECURITY_ERROR` | `terminal` | Authentication or authorization failed. | +| `PROTOCOL_VIOLATION` | `terminal` (forced) | Protocol failure; stop and report it. | + +The Node.js client delivers asynchronous rejections to `onSenderError`; its +default handler logs them. Other clients may use a bounded error inbox that +drops the oldest notification on overflow. Check the client-specific error +handling before relying on a policy override: the Node.js `on_*_error` keys +are currently accepted but do not change these defaults. ## Next steps diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 3ea50b925d..21d230bb26 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -72,11 +72,12 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| | `error_inbox_capacity` | int (≥16) | `256` | Bounded SPSC queue capacity for async error notifications. Overflow drops the oldest entry and increments `getDroppedErrorNotifications`. | -| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | Override the default policy (`HALT` or `DROP_AND_CONTINUE`) for a category. Reserved in the spec but not yet recognised by the connect-string parser. | +| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | All clients accept these keys, but Node.js and Java currently ignore them; .NET applies them. There is no `DROP_AND_CONTINUE` policy. See [Error handling](/docs/connect/clients/connect-string/#error-handling). | -The per-category defaults are documented in +The Node.js defaults are documented in [Concepts § Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). -`PROTOCOL_VIOLATION` and `UNKNOWN` are forced `HALT` and not user-overridable. +`PROTOCOL_VIOLATION` is always terminal; Node.js treats an unknown server +status as retriable rather than silently dropping the batch. ## Other relevant keys @@ -96,8 +97,8 @@ canonical entries. | `auto_flush_rows` | int / `off` | `1000` | Row-count flush trigger. | | `auto_flush_bytes` | int / `off` | `0` (off) | Byte-size flush trigger. | | `auto_flush_interval` | int (ms) / `off` | `100` | Time-since-first-row flush trigger (Node.js: since last flush or sender creation). | -| `init_buf_size` | size | `64K` | Initial encode buffer capacity. | -| `max_buf_size` | size | `100M` | Max encode buffer capacity. | +| `init_buf_size` | size | `64K` | Initial encode buffer capacity; not supported by the Node.js QWP `ws`/`wss` client. | +| `max_buf_size` | size | `100M` | Max encode buffer capacity; not supported by the Node.js QWP `ws`/`wss` client. | | `max_name_len` | int | `127` | Local validation cap for table / column names. | ## Validation From 8a864af0d4ca1ffd2f503cff2a359f69c83ee826 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 19:03:00 +0100 Subject: [PATCH 09/25] docs(nodejs): fix ingestion and recovery contracts --- documentation/changelog.mdx | 2 +- .../connect/clients/connect-string.md | 16 ++- documentation/connect/clients/nodejs.md | 117 ++++++++++++------ .../wire-protocols/qwp-client-behavior.md | 36 ++++-- .../wire-protocols/qwp-ingress-websocket.md | 7 +- .../client-failover/concepts.md | 25 ++-- .../store-and-forward/concepts.md | 12 +- .../store-and-forward/operating-and-tuning.md | 19 ++- 8 files changed, 163 insertions(+), 71 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 711b267b90..cb0c4d3035 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -30,7 +30,7 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, retry policies, HTTP transaction boundaries, replay deduplication, and stale-lock recovery - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index f084e59686..14854b8905 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -660,8 +660,9 @@ Node.js client is the exception: in memory mode, it gives up after - `on` (aliases `sync`, `true`) — retry synchronously on the user thread, up to `reconnect_max_duration_millis`. - `async` — return the `Sender` immediately; the I/O thread retries in - the background indefinitely, surfacing only genuine terminal failures - (auth reject, durable-ack mismatch) via the error inbox. + the background indefinitely, surfacing terminal failures via the error + inbox. Initial authentication rejection remains terminal; see the Node.js + recovery exception below. **Implicit promotion.** Setting any explicit `reconnect_*` key without also choosing an `initial_connect_retry` mode promotes @@ -679,9 +680,14 @@ Node.js client is the exception: in memory mode, it gives up after This is the shutdown data-loss window. Setting it to `0` skips the drain entirely and drops un-ACKed batches on every clean shutdown. -Auth failures during reconnect (authentication rejected, version mismatch, -durable-ack mismatch, non-101 upgrade without a role hint) are immediately -terminal — the loop does not retry them. +Authentication rejection (HTTP `401` / `403`) normally stops the reconnect +loop without trying other hosts. The Node.js client makes an exception after +a regular sender's first successful connection: senders with `sf_dir` or +background memory replay (`initial_connect_retry=async` or `lazy_connect=on`) +retry authentication rejections indefinitely. Initial authentication rejection +remains terminal. This exception does not apply to ordinary memory-only senders +or orphan drainers. See +[authentication during failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). ### Egress failover {#egress-failover} diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 9ac957b6ab..e810634e23 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -735,8 +735,9 @@ state, so borrow one sender per producer (see [Concurrency](#concurrency)). [standalone `Sender`](#standalone-sender). 2. Call `table(name)` to start a row. 3. Add values with the [column methods](#column-methods), such as - `symbol(name, value)` and `doubleColumn(name, value)`. To store a NULL, pass - `null` or `undefined`, or skip the column (see [Null values](#null-values)). + `symbol(name, value)` and `doubleColumn(name, value)`. For a nullable column, + pass `null` or `undefined`, or skip the column to store NULL (see + [Null values](#null-values) for non-nullable defaults). 4. Close the row with `at(timestamp, unit)` or `atNow()`, and `await` the returned promise. It rejects if an auto-flush triggered by the row fails. 5. Repeat from step 2, and call `flush()` to send staged rows. @@ -821,7 +822,8 @@ Names that differ from what you might expect: - `floatColumn()` and `intColumn()` write 64-bit DOUBLE and LONG. Use `float32Column()` and `int32Column()` for FLOAT and INT. - There is no `nullColumn()` or `setNull()`. Pass `null` or `undefined`, or - skip the column. + skip the column; the stored value depends on the column's + [nullability](#null-values). - Arrays use `arrayColumn()`. `doubleArray()` is a [compiled writer](#compiled-object-row-writers) field, not a sender method. - `geohashColumn()` takes raw bits only. Base-32 geohash text is accepted by a @@ -834,14 +836,25 @@ type. A column's type is fixed by the first value a sender stages for it. Writing a different type to the same column later throws -`column type mismatch for ''`. If the table already exists with a -different column type, QuestDB rejects the batch asynchronously; see -[Ingestion errors](#ingestion-errors). +`column type mismatch for ''`. For an existing table, QuestDB rejects +an incompatible type or value asynchronously; see +[Ingestion errors](#ingestion-errors). Compatible conversions are allowed: +for example, `longColumn("price", 123n)` can write to an existing DOUBLE column. +This does not change the sender's local type-consistency rule. ### Null values -To store NULL, pass `null` or `undefined` to any column method, or leave the -column out of the row. All three have the same effect: +Passing `null` or `undefined` to a column method omits the column, just like +leaving it out of the row. For an existing nullable column, QuestDB stores SQL +NULL. BOOLEAN, BYTE, and SHORT are not nullable: omitted BOOLEAN values become +`false`, and omitted BYTE and SHORT values become `0`. + +CHAR uses the zero character as its NULL marker. Current QWP query results can +return that marker as the one-character string `"\u0000"`, rather than JavaScript +`null`. See the [data types](/docs/query/datatypes/overview/) and +[type nullability](/docs/query/datatypes/overview/#type-nullability) references. + +For example, omitting a value for an existing SYMBOL column stores NULL: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -869,16 +882,17 @@ try { - An omitted column is not created on a table that lacks it: a NULL carries no type to infer from. - The column name is still validated when the value is nullish. -- Rows that already exist in a batch, or rows added later, get NULL for any - column they do not set. +- Rows that already exist in a batch, or rows added later, use the same + NULL or non-nullable default for any column they do not set. - INT, LONG, and DATE reserve their minimum values as NULL: writing `-2147483648` to INT or `-9223372036854775808n` to LONG or DATE stores NULL. IPv4 reserves `0.0.0.0` for NULL too, but `ipv4Column()` rejects it with a `RangeError` and discards the row: pass `null` to store an IPv4 NULL. -- A row where every value is nullish is still sent over WebSocket and stored - with NULL in every column. To drop such a row instead, call `cancelRow()` - before closing it. Over UDP, `atNow()` rejects such a row while the sender - knows no non-null column for the table. +- A row where every column value is nullish is still sent over WebSocket. + Its non-designated columns use the NULL/default rules above; the designated + timestamp comes from `at()` or `atNow()`. To drop such a row instead, call + `cancelRow()` before closing it. Over UDP, `atNow()` rejects such a row while + the sender knows no non-null column for the table. ### Designated timestamp @@ -1084,7 +1098,7 @@ try { timestamp: Date.now(), }); - // Arrays, iterables, and async iterables. Absent fields store NULL. + // Arrays, iterables, and async iterables. Absent nullable fields store NULL. await trades.rows([ { symbol: "BTC-USD", side: "buy", price: 39269.98, timestamp: Date.now() }, ]); @@ -1450,10 +1464,15 @@ the directory's lock (see Lock recovery below). its process ID is no longer in use. Otherwise the new sender fails with `QwpReplayStoreLockedError`. This is common in containers: the application usually runs as process ID 1, which is in use again after a restart, and a - replacement container usually has a different host name. Once you are sure - the previous process has exited, delete `.lock.owner` from the journal - directory (`/`, or `/-` for a pooled - sender) and start the sender again. + replacement container usually has a different host name. Once you have + verified that the previous owner has exited and no process is using the slot, + remove its stale `//.lock.owner` directory and restart the sender. + Here `` is ``, or `-` for a pooled sender. If + startup still reports `QwpReplayStoreLockedError`, also inspect + `/.slot-locks/.lock.owner`: this short-lived guard can survive a + crash during lock acquisition or quarantine. Remove that specific owner + directory only after verifying its owner has exited. Never delete the shared + `.slot-locks` directory or another slot's locks. - **Rejected batches.** A batch that QuestDB rejects terminally, such as one with a value of the wrong type for an existing column, stays in the journal. Every new sender on that directory, including the pool's replacement for a @@ -1648,7 +1667,12 @@ Values arrive as these JavaScript types: | GEOHASH | `{ bits: bigint, precisionBits: number }` | | DECIMAL with a precision of 10 or more | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | | DOUBLE[], DOUBLE[][], ... | `{ dimensions: number[], values: number[] }` with values in row-major order | -| NULL of any type | `null` | +| NULL in nullable types other than CHAR | `null` | + +BOOLEAN, BYTE, and SHORT are non-nullable, so omitted values read back as +`false`, `0`, and `0`. A CHAR NULL marker can currently read back as the +one-character string `"\u0000"`, not JavaScript `null`. See +[Null values](#null-values). Some column types cannot be returned over QWP. The server rejects such a query with status `0x06` and a message such as `unsupported column type INTERVAL`. @@ -2081,11 +2105,15 @@ The default policy follows the category: | `not-writable` | Retriable on another endpoint | The server is a replica or cannot accept writes | | `data-loss` | Abandoned | A corrupt store-and-forward journal was set aside | -A retriable rejection is resent. If the same batch keeps being rejected, after +A retriable rejection is resent. For rejections that count toward the +poison-frame detector, if the same batch keeps being rejected after `max_frame_rejections` (4) attempts spanning at least `poison_min_escalation_window_millis` (5 minutes), the sender stops as for a -terminal error. The six `on_*_error` connect-string keys are accepted but not -applied by this client. +terminal error. The `dictionary-gap`, `unknown`, and `not-writable` categories +are exempt: they reset the poison episode instead of adding a strike. +Retriable rejections of symbol-dictionary catch-up frames are also exempt. +The six `on_*_error` connect-string keys are accepted but not applied by this +client. **After a terminal error**, the sender is permanently failed. `waitForAcknowledged()` for the rejected batch rejects with @@ -2231,12 +2259,18 @@ try { } ``` -An authentication failure (HTTP 401 or 403) ends the connection attempt for -the whole endpoint list, because a credential rejected by one node is wrong for -all of them, and it is not retried. The exception is a store-and-forward sender -that has connected before: it keeps retrying, so that a rotated credential -cannot strand its journal. Endpoints in error messages have any embedded -credentials removed. +An authentication rejection (HTTP 401 or 403) is terminal before a sender's +first successful connection and for query connections. It stops the endpoint +walk because credentials are assumed to be shared across the cluster. + +After a successful connection, regular senders with `sf_dir` or background +memory replay (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying +authentication rejections indefinitely. This lets buffered data drain once +server-side authentication is restored. Memory-only senders without those +settings, and orphan drainers, do not have this exception. See +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). + +Endpoints in error messages have any embedded credentials removed. ## Failover and high availability @@ -2425,9 +2459,13 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { await db.close(); ``` -Supplying `egressSession.reconnect`, like setting any `failover*` key, also -makes opening a query connection retry within the failover budget, instead of -failing on the first error. +Supplying an `egressSession.reconnect` options object also makes opening a +query connection retry within the failover budget when the error is retryable. +So does +explicitly setting `failover=on`, or setting a `failover_*` tuning key without +`failover=off`. By contrast, `failover=off` or `egressSession.reconnect: false` +disables the reconnect wrapper. As with other session options, an explicit +`egressSession.reconnect` value replaces the connect-string policy. | Kind | Meaning | |---|---| @@ -2552,9 +2590,10 @@ Version 5.0.0 keeps the ILP API and adds QWP. Changes that affect existing ILP code: - **Null values.** Passing `null` or `undefined` to a column or symbol method now - omits the column, which QuestDB stores as NULL. Earlier versions threw a type - error for most such values. Validate data before calling the sender if you - relied on the error. + omits the column. Existing nullable columns store NULL; BOOLEAN defaults to + `false`, and BYTE and SHORT default to `0` (see [Null values](#null-values)). + Earlier versions threw a type error for most such values. Validate data + before calling the sender if you relied on the error. - **Decimal scale.** `decimalColumn()` over ILP rejects a non-integer `scale` with a `RangeError`. Earlier versions silently coerced it, writing `2.5` as scale 2 and `NaN` as scale 0. @@ -2595,8 +2634,12 @@ try { require `await sender.connect()`. `token=...` selects bearer authentication over HTTP. Over TCP, `username` and `token` set the JWK key ID and private key. -- `flush()` sends the buffer as one HTTP request, which QuestDB commits as one - transaction, and throws if QuestDB rejects it. +- Over HTTP, `flush()` sends the buffer as one request and throws if QuestDB + rejects it. Data is transactional only for a single-table request. A + multi-table request can commit earlier tables before a later table fails, + so a failed flush does not mean no data was committed. Schema changes, such + as automatically added columns, are not rolled back even for a single-table + request. See [HTTP transaction semantics](/docs/connect/compatibility/ilp/overview/#http-transaction-semantics). - Decimals need ILP protocol version 3: HTTP negotiates it automatically, and TCP needs `protocol_version=3`. Arrays need version 2 or later. - Undici is the default HTTP agent. Set `stdlib_http=on` to use the Node.js diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index 2eec7f73fc..29bdd1cde2 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -108,9 +108,10 @@ Why each line matters: - `target=replica` is required to avoid binding a primary/standalone server. The default `target=any` will accept any role. -- `failover=on` is the default. It does **not** affect startup; it only governs - reconnect+replay after a query connection that was already established later - fails during `execute()`. +- `failover=on` is the default. In the Java reference client it does **not** + affect startup; it only governs reconnect+replay after an established query + connection fails during `execute()`. Node.js also uses explicit failover + settings for initial query retries; see the [mental model](#mental-model). --- @@ -127,10 +128,23 @@ share a startup model. You must hold all three in mind: | Query client initial connect | (no mode; always synchronous) | always blocking | | Facade prewarm (how many of each connect at `build()`) | `sender_pool_min`, `query_pool_min` | eager if `min>0`, lazy if `min=0` | -`failover=on` (query default) is **not** a startup setting — it only affects -query execution after a connection exists. This naming trips people up +In Java, `failover=on` (query default) is **not** a startup setting: it only +affects query execution after a connection exists ([sharp edge #3](#known-sharp-edges)). +:::note Node.js query startup + +Node.js also retries initial query connections when `failover=on` is explicitly +set, a `failover_*` tuning key is supplied without `failover=off`, or an +`egressSession.reconnect` options object is supplied. Retries apply only to +retryable failures and stay within the failover budget. With no explicit policy, +failover is enabled for established query connections but initial connection +attempts are not retried. `failover=off` disables the reconnect wrapper; an +explicit `egressSession.reconnect` value overrides the connect-string policy. +See [Node.js connection events](/docs/connect/clients/nodejs/#connection-events). + +::: + ### Ingest initial-connect modes | `initial_connect_retry` | Mode | `build()` behavior on a down server | @@ -283,9 +297,9 @@ here. | --- | --- | --- | | 1 | `initial_connect_retry` is implicitly promoted to `SYNC` when any `reconnect_*` knob is set — a resilience knob silently makes startup block. | Candidate | | 2 | `reconnect_max_duration_millis` is named as if it governs reconnection, but a running sender never consults it — it bounds only the blocking initial connect. | Candidate (naming) | -| 3 | `failover` sounds like it covers startup but only affects post-connect query `execute()`. Queries have no async/lazy initial connect at all. | Candidate | +| 3 | In Java, `failover` sounds like it covers startup but only affects post-connect query `execute()`. Java queries have no async/lazy initial connect. For the Node.js exception, see the [mental model](#mental-model). | Candidate | | 4 | No first-class write-only facade: a write-only user must still supply a query config and remember `query_pool_min=0`, or use `lazy_connect=true`. | Candidate | -| 5 | A single endpoint returning `401`/`403` is treated as cluster-wide terminal and aborts the whole endpoint walk, even at startup, even if other endpoints would accept the credentials. | Intended (documented), revisit | +| 5 | A single endpoint returning `401`/`403` normally aborts the whole endpoint walk, even if other endpoints would accept the credentials. See the [Node.js authentication recovery exception](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). | Intended (documented), revisit | | 6 | Query `serverInfoTimeoutMs` has no config key, so a facade query client cannot tune it. | Candidate | | 7 | The simplest API (`fromConfig` + async) has the worst error visibility — terminal async failures surface only on later producer calls or at `close()`. | Candidate | | 8 | `SenderProgressHandler` has no builder setter on either surface; it must be installed post-construction via `QwpWebSocketSender.setProgressHandler`. | Candidate | @@ -351,6 +365,12 @@ A sender in async mode does not give up because time passed. What ends it is a mismatch, or the poison-frame detector — or the producer hitting `sf_max_total_bytes` and exhausting `sf_append_deadline_millis` on `append()`. +Node.js differs on authentication recovery: initial rejection is terminal, but +regular senders with `sf_dir` or background memory replay +(`initial_connect_retry=async` or `lazy_connect=on`) retry authentication +rejections after their first successful connection. See +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). + ### Reconnect and outage handling **A running sender retries a transport outage indefinitely.** There is no @@ -393,7 +413,7 @@ and apply back-pressure to the producer rather than dropping data. | TLS session/certificate failure | transport error; try next endpoint | | HTTP upgrade timeout / non-auth transport error | try next endpoint | | `421` with `X-QuestDB-Role: REPLICA` | role reject; try next endpoint | -| `401` / `403` auth failure | **terminal**; do not try later endpoints ⚠ | +| `401` / `403` auth failure | **terminal**; do not try later endpoints, except during [Node.js background sender recovery](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | | durable-ack requested but unsupported | terminal mismatch | | successful write upgrade | bind this endpoint | | all endpoints fail transport | throw / retry per initial/reconnect mode | diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index e651e7a849..55c9d78f3f 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1161,8 +1161,11 @@ Key behaviors: The Node.js client is the exception: it applies `zone=` and `target=` to ingress too, so `target=replica` in a shared connect string stops its ingestion. -- **Authentication errors are terminal** at any host (`401`/`403`). The - reconnect loop does not continue past them. +- **Authentication rejection (`401`/`403`) is normally terminal.** After a + successful connection, a regular Node.js sender with `sf_dir` or background + memory replay (`initial_connect_retry=async` or `lazy_connect=on`) instead + retries it indefinitely. Initial authentication rejection remains terminal; + see the [authentication recovery exception](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). - **`421 + X-QuestDB-Role`** is a role reject: transient if the role is `PRIMARY_CATCHUP`, topology-level otherwise. - **All other upgrade errors are transient** and feed into the reconnect loop, diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index a5519df0ef..85650a8998 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -191,7 +191,7 @@ will not help. | Condition | Why terminal | |---|---| -| HTTP `401` / `403` on upgrade | Credentials are cluster-wide; retrying floods server logs without recovery. | +| HTTP `401` / `403` on upgrade | Credentials are assumed to be cluster-wide. See the [Node.js sender recovery exception](#authentication-is-cluster-wide). | | Server-status reject (SF) | Application-layer reject; replay reproduces the same response. | ### Topology — handled inside the round @@ -237,14 +237,21 @@ when at least one peer is healthy." ## Authentication is cluster-wide -A `401` or `403` on the HTTP upgrade is terminal — the client does not retry -other hosts. The assumption is that auth credentials are configured -identically across the cluster, so a credential failure against one node is -a credential failure against all of them. Retrying would spam every peer's -audit log without recovering. - -If your deployment has per-host credentials, that is unsupported and outside -the failover model — split the workload into one connect string per credential. +A `401` or `403` on the HTTP upgrade is normally terminal: the client does not +retry other hosts. Credentials are assumed to be configured identically across +the cluster, so trying another node would repeat the rejection. + +The Node.js client makes an exception for a regular ingress sender that has +already connected successfully and uses `sf_dir` or background memory replay +(`initial_connect_retry=async` or `lazy_connect=on`). It retries authentication +rejections indefinitely so buffered data can drain after server-side +authentication is restored. Initial authentication rejection remains terminal, +including in these modes. The exception does not apply to query connections, +ordinary memory-only senders, or orphan drainers. See +[Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors). + +Per-host credentials are outside the failover model. Use a separate connect +string for each credential. ## Next steps diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 192c0e5cc3..169e84768b 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -65,8 +65,9 @@ Two distinct counters track frame identity: - **FSN** (frame-sequence-number) — a monotonic counter assigned when a frame is appended to the substrate. FSN survives reconnects and (in SF mode) restarts. It is the substrate's permanent identifier for a frame. -- **wireSeq** — the per-connection counter the server uses for - deduplication, reset to `0` on every successful WebSocket upgrade. +- **wireSeq** is the per-connection counter used to correlate acknowledgements + with sent frames. It resets to `0` on every successful WebSocket upgrade + and is not a row-deduplication key. On every (re)connect the relationship is pinned: @@ -160,8 +161,11 @@ On every successful (re)connect: 2. `wireSeq` resets to `0`. 3. The read cursor rewinds to the first un-acked frame on disk (or in memory). -4. Frames stream to the wire in FSN order. The server's dedup window - absorbs any frames that landed before the disconnect. +4. Frames stream to the wire in FSN order. A frame committed before the + disconnect but not acknowledged can be inserted again: replay is at least + once. Prevent duplicate rows with table-level `DEDUP UPSERT KEYS` covering + the designated timestamp and stable source identity, preserving both on + retries. See [Deduplication](/docs/concepts/deduplication/). 5. New frames appended by the producer during replay are picked up automatically — the I/O loop watches a volatile `publishedFsn` cursor. diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index d6a24df0a2..85a26347de 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -63,9 +63,18 @@ crashed Node.js sender therefore leaves the slot locked: a new sender takes it over automatically only on the same host, once the recorded process ID is no longer in use. That often fails in containers, where the application usually runs as process ID 1 and a replacement container has a new host name. -Otherwise, make sure the previous process has exited, then delete -`.lock.owner`. Node.js and other clients do not see each other's locks, so -never let them use the same `sf_dir` at the same time. See the +Otherwise, verify that the previous owner has exited and no process is using +that slot before removing its stale `//.lock.owner` directory. +Here `` is ``, or `-` for a pooled Node.js sender. + +If startup still reports `QwpReplayStoreLockedError`, inspect +`/.slot-locks/.lock.owner` too. This short-lived guard can survive +a crash during lock acquisition or quarantine. Remove only that specific +owner directory after verifying its owner has exited, never the shared +`.slot-locks` directory or another slot's locks. + +Node.js and other clients do not see each other's locks, so never let them use +the same `sf_dir` at the same time. See the [Node.js client](/docs/connect/clients/nodejs/#store-and-forward). ::: @@ -108,8 +117,8 @@ when the new one comes up. Solutions: - Stop the previous process. For clients using OS locks, the kernel releases the lock on exit (even after `kill -9`). A killed Node.js sender can leave - `.lock.owner` behind: verify the old owner is gone before removing it; see - [`.lock` and `.lock.pid`](#lock-and-lockpid). + stale owner directories behind: verify the old owner is gone before removing + the specific directories described in [`.lock` and `.lock.pid`](#lock-and-lockpid). - Use a deployment unit that orders shutdown before startup. - For containerised deployments, set `sender_id` from a per-pod stable identity so two pods with the same template name don't collide. From 957fe3f751ea5987870fc7566880375871715471 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 19:56:46 +0100 Subject: [PATCH 10/25] docs(nodejs): address QWP review findings --- documentation/changelog.mdx | 2 +- .../connect/clients/connect-string.md | 9 ++- documentation/connect/clients/nodejs.md | 65 ++++++++++++++----- .../store-and-forward/concepts.md | 27 +++++--- .../store-and-forward/configuration.md | 2 +- .../store-and-forward/operating-and-tuning.md | 24 +++++-- 6 files changed, 95 insertions(+), 34 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index cb0c4d3035..2b3a17e28b 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -30,7 +30,7 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, retry policies, HTTP transaction boundaries, replay deduplication, and stale-lock recovery +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, ordinary versus background-memory retry policies, HTTP transaction boundaries, replay deduplication, stale-lock recovery, journal capacity allowances, configuration override precedence, and reuse of existing PGWire table schemas - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 14854b8905..8bf5b5f2c4 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -543,10 +543,15 @@ equivalent — same architecture, no durability across restarts. - `sf_max_segment_bytes` — per-segment rotation threshold. Must be ≥ the largest single flushed frame. Default: `4 MiB` (`4m`). Accepts [size suffixes](#size-suffixes). -- `sf_max_total_bytes` — hard cap on per-slot storage. When the slot - reaches the cap, `append()` blocks until ACKs trim space (see +- `sf_max_total_bytes` controls per-slot capacity for producer backpressure. + When capacity is exhausted, `append()` blocks until ACKs trim space (see `sf_append_deadline_millis`). Defaults: `10 GiB` (`10g`) in SF mode, `128 MiB` (`128m`) in memory mode. Accepts size suffixes. + On Node.js with `sf_dir`, this is a journal size target, not a hard disk + limit: transaction-closing batches and retained symbol dictionaries can + exceed it, and other metadata needs additional space. Provision disk + headroom; see the [Node.js capacity guidance](/docs/connect/clients/nodejs/#store-and-forward). + Without `sf_dir`, the key caps the in-memory replay queue. ### Sender restart and replay diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index e810634e23..13e0b054ed 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -50,6 +50,8 @@ documents the recommended QWP path. For ILP, see default) at `/write/v4` for ingestion and `/read/v1` for queries. If QuestDB is not running yet, see the [quick start](/docs/getting-started/quick-start/). + + ## Installation ```shell @@ -68,7 +70,17 @@ such as `as const`. ## Quick start Connect with one connect string, write two rows, and try to query the ETH-USD -row. Ingestion is asynchronous, so an immediate read may not see it yet: +row. Ingestion is asynchronous, so an immediate read may not see it yet. + +:::note Existing `trades` tables + +This example assumes `trades` does not exist yet. If it already exists, QWP +uses its existing designated timestamp column. For the `trades(ts, ...)` schema +in the [PGWire guide](/docs/connect/compatibility/pgwire/nodejs/), +replace `SELECT timestamp` with `SELECT ts` in the query below. The sender's +`at()` calls need no change: they write to the existing designated timestamp. + +::: ```typescript import { @@ -145,8 +157,8 @@ What happens: resolves even when that wait times out; see [Closing the pooled client](#closing-the-pooled-client). -The table was created automatically by the first row, so its designated -timestamp column is named `timestamp`. Timestamps come back as `bigint` +If `trades` did not exist, ingestion creates it automatically with a designated +timestamp column named `timestamp`. Timestamps come back as `bigint` microseconds since the Unix epoch; see [Reading result values](#reading-result-values) for every type. @@ -386,8 +398,8 @@ For every key and its default, see the ### Programmatic options Callbacks, custom agents, and other settings a string cannot express go in the -second argument. The connect string is validated in full first; when both set -the same option, the typed value wins: +second argument. When the connect string and typed options set the same option, +the typed value wins: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -429,6 +441,8 @@ from the connect string. When you supply the object, for example to register ::: + + ## Authentication and TLS QWP authenticates on the WebSocket upgrade request, before any data is @@ -726,6 +740,8 @@ acknowledgement before returning each sender (see ## Data ingestion + + ### General usage pattern A sender is not safe for concurrent producers: the row in progress is shared @@ -997,6 +1013,8 @@ with `long arrays are not supported, only double arrays`. Query results return arrays as `{ dimensions, values }`; see [Reading result values](#reading-result-values). + + ### Decimals Create decimal columns ahead of time with the precision you need. QWP can @@ -1041,10 +1059,12 @@ try { } ``` -- `decimalColumnText()` takes a decimal string, scientific notation included - (`"1.5e-3"`), and preserves the literal's scale, including trailing zeros. - Passing a `number` works, but JavaScript drops trailing zeros when formatting. -- `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` +- `decimalColumnText()` takes a + decimal string, scientific notation included (`"1.5e-3"`), and preserves the + literal's scale, including trailing zeros. Passing a `number` works, but + JavaScript drops trailing zeros when formatting. +- + `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` or as big-endian two's-complement bytes in an `Int8Array`. - `decimal64Column()`, `decimal128Column()`, and `decimal256Column()` take an unscaled `bigint` and select the wire width directly. @@ -1433,10 +1453,10 @@ try { ``` With a journal, the sender keeps accepting rows while QuestDB is unreachable, -until `sf_max_total_bytes` (10 GiB) is full. It retries the connection -indefinitely once it has connected, and a new sender opened on the same -directory replays what the previous process left behind, once it can take over -the directory's lock (see Lock recovery below). +subject to journal capacity and backpressure as described below. It retries the +connection indefinitely once it has connected, and a new sender opened on the +same directory replays what the previous process left behind, once it can take +over the directory's lock (see Lock recovery below). - **Layout.** A standalone `Sender` journals into `/`. A pooled client uses one directory per pooled sender: @@ -1449,9 +1469,18 @@ the directory's lock (see Lock recovery below). not a power loss. `periodic` checkpoints in the background every `sf_sync_interval_millis` (5 seconds). `append` makes every append durable before `flush()` resolves. -- **Backpressure.** When the journal is full, publishing waits up to - `sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, - then rejects with `QwpReplayStoreAppendTimeoutError`. +- **Capacity.** With `sf_dir`, `sf_max_total_bytes` (10 GiB by default) is a + journal size target, not a hard disk limit. Transaction-closing batches can + reserve extra segments so a full journal does not block the commit needed + to release space. Segment reservations can reach roughly twice the target, + depending on segment rounding; retained symbol dictionaries and other + metadata take additional space. Provision headroom for every sender and + monitor actual disk usage. Without `sf_dir`, the key caps the in-memory + replay queue instead. +- **Backpressure.** When an append cannot fit within the journal's capacity + allowances, publishing waits up to `sf_append_deadline_millis` (30 seconds) + for acknowledgements to free space, then rejects with + `QwpReplayStoreAppendTimeoutError`. - **Startup.** `lazy_connect=on` lets the pooled client start while QuestDB is down, as in the example above. `initial_connect_retry=async` alone is not enough for the pooled client, because its query pool still connects at @@ -2512,6 +2541,8 @@ Node.js runs your code on one thread, but async functions interleave at every Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy work out of them. + + ## Configuration reference The [connect string reference](/docs/connect/clients/connect-string/) documents @@ -2536,7 +2567,7 @@ every key. The Node.js client's defaults and deviations: | `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | | `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | | `sf_durability` | `memory` | `memory`, `periodic`, or `append`. | -| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | Journal or memory queue cap. | +| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](#store-and-forward), not a hard disk limit; memory queue cap without `sf_dir`. | | `sf_max_segment_bytes` | `4m` | Journal segment size, which also caps a batch. | | `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | | `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. | diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 169e84768b..dd570fabcf 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -38,9 +38,10 @@ SF runs in either of two modes selected by the connect string: Both modes share the same reconnect loop, the same backoff and retry budgets, and the same on-the-wire behaviour. The only difference is -where unacked data lives. The Node.js client is the exception: in memory -mode, a sender without `initial_connect_retry=async` gives up after -`reconnect_max_duration_millis` (5 minutes by default); see the +where unacked data lives. The Node.js client is the exception: ordinary +memory mode gives up after `reconnect_max_duration_millis` (5 minutes by +default). Background memory mode (`initial_connect_retry=async`, also selected +by pooled `lazy_connect=on`) and disk-backed SF retry indefinitely. See the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). ## What "frame" means here @@ -151,9 +152,13 @@ When the wire connection breaks — for any reason — the I/O thread enters the reconnect loop documented in [Client failover concepts](/docs/high-availability/client-failover/concepts/). The producer is **not notified**: it keeps publishing into the substrate, -bounded by `sf_max_total_bytes` (see backpressure below). On the Node.js -client in memory mode, `flush()` instead waits for the reconnect, up to -`reconnect_max_duration_millis`. +subject to available capacity (see [Backpressure](#backpressure)). + +On Node.js, only ordinary memory mode waits for reconnect in `flush()`, up to +`reconnect_max_duration_millis`. Background memory mode, enabled by +`initial_connect_retry=async` or pooled `lazy_connect=on`, keeps accepting +batches into the memory replay queue until capacity is exhausted and retries +indefinitely. See the [three Node.js flushing modes](/docs/connect/clients/nodejs/#flushing). On every successful (re)connect: @@ -175,12 +180,18 @@ in the `getTotalFramesReplayed` observability counter. ## Backpressure -The substrate enforces `sf_max_total_bytes` as a hard cap on resident -storage. When the cap is hit, the producer's `appendBlocking` call +The substrate uses `sf_max_total_bytes` to apply backpressure on resident +storage. When capacity is exhausted, the producer's `appendBlocking` call busy-spins (with cooperative yield) up to `sf_append_deadline_millis` waiting for ACK-driven trim to free space. If the deadline fires, the call throws a typed exception. +On Node.js with `sf_dir`, this value is a journal size target rather than a +hard disk limit. Transaction-closing batches and retained symbol dictionaries +can exceed it, and other metadata needs additional space. Provision headroom +for each sender; see [Node.js journal capacity](/docs/connect/clients/nodejs/#store-and-forward). +Without `sf_dir`, the key caps the in-memory replay queue. + The exception message distinguishes the two scenarios: - **Backpressure while the wire is publishing** — the server is acking diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 21d230bb26..d5faef774e 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -27,7 +27,7 @@ mode. | `sf_dir` | path | unset | Group root directory. When set, the slot lives at `//` and unacked data is durable across process restarts. When unset, the substrate runs in memory mode. | | `sender_id` | string | `default` | Slot subdirectory name. Two senders sharing the same `sender_id` and `sf_dir` will collide on the slot lock. Must not contain path separators or be empty. | | `sf_max_segment_bytes` | size | `4M` | Per-segment file size; rotation threshold. | -| `sf_max_total_bytes` | size | `128M` (memory) / `10G` (SF) | Hard cap on resident SF storage. Triggers producer backpressure when full. | +| `sf_max_total_bytes` | size | `128M` (memory) / `10G` (SF) | Capacity for producer backpressure. Node.js disk journals can exceed this target for transaction completion and symbol dictionaries; provision [additional disk headroom](/docs/connect/clients/nodejs/#store-and-forward). Without `sf_dir`, this caps the memory replay queue. | | `sf_durability` | enum | `memory` | `memory` relies on the page cache; `periodic` checkpoints in the background and requires `sf_dir`. Node.js also supports `append`, which makes each journal append durable before `flush()` resolves. Go and .NET accept only `memory`; other clients reject `append` at build time. `flush` is not supported. | | `sf_sync_interval_millis` | int (ms) | `5000` | Checkpoint cadence for `sf_durability=periodic`; rejected without it. A floor, not a guarantee: scheduler and storage latency add to it. | | `sf_append_deadline_millis` | int (ms) | `30000` | How long a producer `appendBlocking` call waits for ACK-driven trim to free space before throwing. | diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index 85a26347de..1810f292b3 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -144,10 +144,22 @@ disk usage staying high under slow ack cadence. ### `sf_max_total_bytes` — slot capacity (default `128 MiB` memory / `10 GiB` SF) -This is the **hard cap** on resident SF storage — sealed segments plus -the active segment. When this fills, producer `appendBlocking` calls -block (with cooperative yield) for up to `sf_append_deadline_millis` -waiting for ACK-driven trim to free space; on timeout the call throws. +This controls resident SF storage: sealed segments plus the active segment. +When capacity is exhausted, producer `appendBlocking` calls block (with +cooperative yield) for up to `sf_append_deadline_millis` waiting for ACK-driven +trim to free space; on timeout the call throws. + +:::caution Node.js disk-journal capacity + +With `sf_dir`, Node.js treats `sf_max_total_bytes` as a target, not a hard disk +limit. Transaction-closing batches can reserve extra segments to make the +commit possible when the journal is full. Segment reservations can reach +roughly twice the target, depending on segment rounding, with retained symbol +dictionaries and other metadata requiring additional space. Provision +headroom per sender and monitor actual disk usage; do not use the target as a +filesystem quota. See [Node.js store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). + +::: Size this against your **worst expected outage** times your ingest rate: @@ -320,7 +332,9 @@ When several senders share a host and a `sf_dir`: sender. - Consider `drain_orphans=on` if dynamic sender identities mean dead instances can leave permanent orphans. -- Size `sf_max_total_bytes × number_of_senders` against available disk. +- Size `sf_max_total_bytes × number_of_senders` against available disk. For + Node.js, also reserve each sender's transaction and metadata headroom (see + [Sizing capacity](#sizing-capacity)). - Plan for the worst-case lock-collision recovery: a misconfigured fleet that all share `sender_id=default` will leave only one sender alive on each host. That is the design — fail loudly rather than From c54acacccffe6d6ec1cde3183ef7769024a4bf80 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Wed, 30 Sep 2026 22:45:14 +0100 Subject: [PATCH 11/25] docs(nodejs): clarify QWP errors, timeouts, and retry behavior --- documentation/changelog.mdx | 2 +- .../clients/date-to-timestamp-conversion.md | 2 + documentation/connect/clients/nodejs.md | 165 +++++++++++------- .../wire-protocols/qwp-client-behavior.md | 18 +- .../store-and-forward/concepts.md | 10 +- .../store-and-forward/configuration.md | 6 +- .../store-and-forward/when-to-use.md | 19 +- 7 files changed, 141 insertions(+), 81 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 2b3a17e28b..e42ca7e273 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -30,7 +30,7 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, ordinary versus background-memory retry policies, HTTP transaction boundaries, replay deduplication, stale-lock recovery, journal capacity allowances, configuration override precedence, and reuse of existing PGWire table schemas +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, ordinary versus background-memory retry policies, HTTP transaction boundaries, replay deduplication, stale-lock recovery, journal capacity allowances, configuration override precedence, and reuse of existing PGWire table schemas. Also clarified cancellation and handshake deadlines, typed pool settings, row validation versus publication failures, acknowledgement error timing, and background durable-ack retries - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/date-to-timestamp-conversion.md b/documentation/connect/clients/date-to-timestamp-conversion.md index 0970224ed3..4fc51b5d63 100644 --- a/documentation/connect/clients/date-to-timestamp-conversion.md +++ b/documentation/connect/clients/date-to-timestamp-conversion.md @@ -319,6 +319,8 @@ class Program ``` Learn more about the [QuestDB .NET Client](/docs/connect/clients/dotnet/) + + ## Date to Timestamp in JavaScript/Node.js A JavaScript `Date` stores milliseconds since the Unix epoch. The QuestDB diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 13e0b054ed..28f606bbe6 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -413,7 +413,11 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { console.error("rejected batch", error.category, error.serverMessage), }, // Query session defaults - egressSession: { queryTimeoutMs: 30_000 }, + egressSession: { + queryTimeoutMs: 30_000, + cancelDrainTimeoutMs: 5_000, + serverInfoTimeoutMs: 10_000, + }, // Egress-only routing and compression egress: { compression: "zstd" }, // Pool sizes and timeouts @@ -496,14 +500,21 @@ To route the connection through an HTTP or SOCKS proxy, pass an agent such as `Sender`). A custom agent owns certificate verification, so it cannot be combined with `tls_verify` or `tls_roots`. -Two deadlines bound connection setup: `connect_timeout` covers DNS and the -TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and +Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS +and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits -`connect_timeout` when only the latter is set. A timeout produces a -`QwpUpgradeError` whose `timeoutPhase` is `connect` or `authentication`. The -pooled client reports it, like every failure to open a connection, as the -`cause` of a `QwpPoolResourceError`; see -[Connection-level errors](#connection-level-errors). +`connect_timeout` when only the latter is set. A timeout in either phase +produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or +`authentication`. + +After the upgrade, a query connection has a separate 5-second deadline for +the initial QWP `SERVER_INFO` frame. Configure it with the typed option +`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not +change it. Expiry produces an ordinary `Error` with the message +`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. + +The pooled client reports connection setup failures as the `cause` of a +`QwpPoolResourceError`; see [Connection-level errors](#connection-level-errors). ### Unsupported authentication paths @@ -578,7 +589,7 @@ try { .at(Date.now(), "ms"); } } finally { - // Flushes the rows and returns the sender. Does not wait for the ACK. + // With default options, flushes and returns without waiting for ACKs. await sender.close(); } } finally { @@ -594,8 +605,10 @@ that hold a sender at the same time. `close()` on a borrowed sender flushes completed rows, discards an unfinished row with a warning, and returns the sender to the pool. It does not close the -WebSocket and does not wait for QuestDB to acknowledge the rows. To confirm -delivery before returning the sender, call `flush()` and then +WebSocket or wait for acknowledgements by default. With `awaitServerAck: true` +or `awaitDurableAck: true`, the flush performed by `close()` waits for its +acknowledgement too. To confirm delivery of all previously published rows +before returning the sender, call `flush()` and then `waitForAcknowledged(sender.publishedSequence)`; see [Awaiting acknowledgements](#awaiting-acknowledgements). @@ -664,8 +677,17 @@ Starting a second query on a lease while one is still active throws | `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | | `lazy_connect` | `off` | Start without connecting. See below. | -The typed equivalents live in the `pool` section of the second argument +Pool sizes, acquisition and idle timeouts, lifetime, and housekeeping settings +have typed equivalents in the `pool` section of the second argument (`senderPoolMin`, `acquireTimeoutMs`, `housekeepingIntervalMs`, and so on). +The other two settings use different locations: + +- `query_close_timeout_ms` maps to `egressSession.cancelDrainTimeoutMs`, not + `pool`. +- Set `lazy_connect=on` in the connect string. When passing a full + `QwpNodeClientOptions` object instead of a string, use top-level + `lazyConnect: true`. It is not supported in `pool` or the second argument. + When creating a new pooled connection fails, the borrow rejects with `QwpPoolResourceError`, whose `cause` holds the connection error. @@ -766,19 +788,13 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); try { const sender = await db.borrowSender(); try { - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.25) - .at(Date.now(), "ms"); - } catch (error) { - // An invalid value throws a TypeError or RangeError. The row in progress - // is discarded; rows completed earlier stay staged. - console.error("row rejected:", error); - } + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.25) + .at(Date.now(), "ms"); await sender.flush(); } finally { await sender.close(); @@ -793,13 +809,19 @@ below. Table and column names are validated locally with QuestDB's rules (at most 127 UTF-8 bytes by default, see `max_name_len`), and column names are case-insensitive: the first spelling used is kept. -When a column method or `at()` rejects a value, the sender discards the whole -row in progress, including its table, so a half-built row never reaches -QuestDB. The next row must start with `table()` again; a column method called -before that throws `table name must be set before adding columns`. +When local value validation in a column method or `at()` fails, the sender +discards the whole row in progress, including its table, so a half-built row +never reaches QuestDB. The next row must start with `table()` again; a column +method called before that throws `table name must be set before adding columns`. `cancelRow()` discards a row in progress without an error, and `reset()` also drops every row staged since the last flush. +An awaited `at()` or `atNow()` can also reject because an auto-flush failed +after the row was completed. This does not mean the row was discarded: +completed rows can remain staged for a later `flush()` or `close()`, or be +queued for replay. Do not blindly resubmit a row because its `at()` promise +rejected. See [Flushing](#flushing) and [Ingestion errors](#ingestion-errors). + ### Column methods These methods are available on pooled senders and on senders from @@ -851,12 +873,19 @@ The standalone `Sender` class exposes only `symbol`, `stringColumn`, type. A column's type is fixed by the first value a sender stages for it. Writing a -different type to the same column later throws -`column type mismatch for ''`. For an existing table, QuestDB rejects -an incompatible type or value asynchronously; see -[Ingestion errors](#ingestion-errors). Compatible conversions are allowed: -for example, `longColumn("price", 123n)` can write to an existing DOUBLE column. -This does not change the sender's local type-consistency rule. +different type to the same column in a later row throws +`column type mismatch for ''`. + +Within one row, duplicate column assignments keep the first value, including +names that differ only in case. For example, +`.doubleColumn("price", 1).stringColumn("PRICE", "wrong")` keeps `1` and does +not raise a type mismatch. Invalid values can still fail local validation. + +For an existing table, QuestDB rejects an incompatible type or value +asynchronously; see [Ingestion errors](#ingestion-errors). Compatible +conversions are allowed: for example, `longColumn("price", 123n)` can write to +an existing DOUBLE column. This does not change the sender's local +type-consistency rule. ### Null values @@ -1223,7 +1252,7 @@ flush, then write the other rows again. up to `close_flush_timeout_millis` (5 seconds by default) for their acknowledgement. `0` or a negative value skips the wait. An unfinished row is discarded with a warning. On a borrowed sender, `close()` flushes and returns -the sender to the pool without waiting; see +the sender to the pool without waiting for acknowledgements by default; see [Borrowing a sender](#borrowing-a-sender). If the acknowledgement does not arrive in time, `close()` on a standalone @@ -1328,7 +1357,9 @@ wait for every row written so far, call `flush()` and wait for To make every `flush()` wait for its acknowledgement, set `awaitServerAck`: `connectQwpNodeClient(conf, { sender: { awaitServerAck: true } })`, or `{ qwp: { sender: { awaitServerAck: true } } }` for a standalone `Sender`. A -server rejection then rejects `flush()` itself, with `QwpIngressNackError`. +server rejection then rejects the waiting `flush()` itself. See +[Ingestion errors](#ingestion-errors) for the error classes before and after +a terminal failure. Acknowledgement is not required for delivery: unacknowledged batches are replayed after a reconnect, and a standalone sender waits for them on @@ -1564,10 +1595,14 @@ wss::addr=db.example.com:9000;token=YOUR_TOKEN;request_durable_ack=on; To make every flush wait for durability, add the typed option `{ sender: { awaitDurableAck: true } }`. If the server does not support durable acknowledgement, connecting fails with `QwpDurableAckUnavailableError`, which -the pooled client reports as the `cause` of a `QwpPoolResourceError`. A sender -that connects in the background, with `initial_connect_retry=async` or -`lazy_connect=on`, does not fail: it keeps retrying and emits -`durable-ack-unavailable` [connection events](#connection-events). +the pooled client reports as the `cause` of a `QwpPoolResourceError`. + +Background senders with `initial_connect_retry=async` or `lazy_connect=on` +instead keep retrying and emit `durable-ack-unavailable` +[connection events](#connection-events). A store-and-forward sender does the +same when reconnecting after its first successful connection. Monitor these +events and buffer usage: successful background startup does not confirm that +the server supports durable acknowledgement. ### Fire-and-forget UDP @@ -1875,10 +1910,11 @@ A query ends early in four ways: - **Deadline.** Set a default with `egressSession: { queryTimeoutMs }`, or per query with `timeoutMs`. On expiry, iteration and `completion` reject with `QwpEgressQueryTimeoutError` and the client sends a cancel to QuestDB. -- **Cancel.** `await query.cancel()` asks QuestDB to stop. Iteration and +- **Cancel.** `await query.cancel()` sends a request to stop; it does not wait + for QuestDB to stop. When QuestDB processes the cancel, iteration and `completion` reject with `QwpEgressQueryError` whose `status` is `0x0a` - (CANCELLED). A query that has already finished, or a DDL or DML statement, - completes normally instead. + (CANCELLED). A query that finishes before the cancel is processed, or a DDL + or DML statement, completes normally instead. - **Leaving the loop.** `break`, `return`, or an exception inside `for await` cancels the query, and `completion` rejects with `QwpEgressQueryAbandonedError`. @@ -1888,19 +1924,21 @@ A query ends early in four ways: QuestDB acts on a cancel between result batches, but while a query streams without a [credit window](#flow-control), it may not read the cancel until the -whole result is sent. Without a credit window: - -- `cancel()` can end with `QwpEgressQueryCancelTimeoutError` instead of the - `0x0a` rejection: the client waits `query_close_timeout_ms` (5 seconds) for - QuestDB to stop, then closes the connection. +whole result is sent. Set `initial_credit` in the connect string or +`initialCredit` per query to improve cancellation responsiveness while +streaming. For example, start with 1 MiB; the client replenishes it as your +loop consumes batches. Credit does not interrupt expensive work before the +next batch or bound cancellation latency. + +With or without a credit window: + +- After an explicit cancel, iteration and `completion` can reject with + `QwpEgressQueryCancelTimeoutError` instead of the `0x0a` rejection: the + client waits `query_close_timeout_ms` (5 seconds) for QuestDB to stop, then + closes the connection. - After a deadline or an early exit from the loop, returning the lease can take up to twice `query_close_timeout_ms`. -Set a credit window when you cancel queries or use deadlines, with -`initial_credit` in the connect string or `initialCredit` per query. About -1 MiB is enough: the client replenishes it as your loop consumes batches, and -QuestDB then stops within milliseconds. - ```typescript import { connectQwpNodeClient, @@ -1913,7 +1951,7 @@ try { try { const query = await lease.query( "SELECT symbol, avg(price) FROM trades SAMPLE BY 1m", - // The credit window lets QuestDB act on the cancel promptly. + // Credit limits streaming ahead, not cancellation latency. { timeoutMs: 5_000, initialCredit: 1024 * 1024 }, ); for await (const batch of query) { @@ -2144,13 +2182,16 @@ Retriable rejections of symbol-dictionary catch-up frames are also exempt. The six `on_*_error` connect-string keys are accepted but not applied by this client. -**After a terminal error**, the sender is permanently failed. -`waitForAcknowledged()` for the rejected batch rejects with -`QwpIngressNackError`, and every later `flush()` or `close()` rejects with -`QwpReplayRejectedError`, whose `status` and message repeat the server's. Close -the sender and create a new one. A pooled sender is replaced automatically -after the `close()` that reports the error. What happens to the rejected batch -depends on the mode: +**After a terminal server rejection**, the sender is permanently failed. An +already-pending `waitForAcknowledged()` for the rejected batch can reject with +`QwpIngressNackError`. Once the terminal failure is latched, new calls to +`waitForAcknowledged()`, `flush()`, or `close()` reject with +`QwpReplayRejectedError`, whose `status` and message repeat the server's. +Error handlers must allow either class depending on timing. + +Close the sender and create a new one. A pooled sender is replaced +automatically after the `close()` that reports the error. What happens to the +rejected batch depends on the mode: - **Without store-and-forward**, the failed sender's unacknowledged batches, including the rejected one, are discarded with it, and the new sender starts diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index 29bdd1cde2..7a02721ac4 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -360,10 +360,11 @@ With `initial_connect_retry=async`: - Terminal errors go to a configured `SenderErrorHandler`; without one they surface on later producer calls or at close-time. -A sender in async mode does not give up because time passed. What ends it is a -**terminal** condition — authentication rejection, a durable-ack capability -mismatch, or the poison-frame detector — or the producer hitting -`sf_max_total_bytes` and exhausting `sf_append_deadline_millis` on `append()`. +A sender in async mode does not give up because time passed. Terminal +conditions include authentication rejection, a durable-ack capability +mismatch, and poison-frame escalation, subject to the Node.js exceptions +below. Producer calls can also fail when the buffer reaches +`sf_max_total_bytes` and exhausts `sf_append_deadline_millis` on `append()`. Node.js differs on authentication recovery: initial rejection is terminal, but regular senders with `sf_dir` or background memory replay @@ -371,6 +372,13 @@ regular senders with `sf_dir` or background memory replay rejections after their first successful connection. See [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). +For unsupported durable acknowledgement, Node.js senders with +`initial_connect_retry=async` or `lazy_connect=on` keep retrying and emit +`durable-ack-unavailable`, rather than becoming terminal. Store-and-forward +senders do the same after their first successful connection. Monitor these +[connection events](/docs/connect/clients/nodejs/#connection-events) and buffer +usage; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). + ### Reconnect and outage handling **A running sender retries a transport outage indefinitely.** There is no @@ -414,7 +422,7 @@ and apply back-pressure to the producer rather than dropping data. | HTTP upgrade timeout / non-auth transport error | try next endpoint | | `421` with `X-QuestDB-Role: REPLICA` | role reject; try next endpoint | | `401` / `403` auth failure | **terminal**; do not try later endpoints, except during [Node.js background sender recovery](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | -| durable-ack requested but unsupported | terminal mismatch | +| durable-ack requested but unsupported | terminal mismatch, except for [Node.js background retries](/docs/connect/clients/nodejs/#durable-acknowledgement) | | successful write upgrade | bind this endpoint | | all endpoints fail transport | throw / retry per initial/reconnect mode | | all endpoints role-reject as replicas | `QwpRoleMismatchException` | diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index dd570fabcf..b636029303 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -132,10 +132,12 @@ object store** (S3, Azure Blob, GCS, or NFS). watermarks. The client matches the head of the OK queue against these watermarks; each fully-covered head entry pops, and `ackedFsn` advances to the highest covered wireSeq. -- The client opt-in is mandatory — the connect fails loudly if the server - does not echo `X-QWP-Durable-Ack: enabled` on the upgrade response. - This avoids the silent failure mode where the producer waits forever - for ack frames that will never arrive. +- The client requires an `X-QWP-Durable-Ack: enabled` echo on the upgrade + response and rejects a connection without it, rather than waiting for ack + frames it cannot receive. This is normally terminal. Node.js background + senders instead retry and emit `durable-ack-unavailable`; see + [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement) + for the retry modes and monitoring requirements. Durable-ack mode is the right choice when "data is in the object store" is the durability bar, but it has two costs: a longer time-to-trim (so diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index d5faef774e..1e238cbdf4 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -64,7 +64,7 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| -| `request_durable_ack` | bool | `off` | Opt-in via the upgrade header `X-QWP-Request-Durable-Ack: true`. Trim is then driven by `STATUS_DURABLE_ACK` frames only; OK frames no longer advance the trim watermark. Connect fails loudly if the server does not echo `X-QWP-Durable-Ack: enabled`. WebSocket transports only. | +| `request_durable_ack` | bool | `off` | Opt-in via the upgrade header `X-QWP-Request-Durable-Ack: true`. Trim is then driven by `STATUS_DURABLE_ACK` frames only; OK frames no longer advance the trim watermark. A missing `X-QWP-Durable-Ack: enabled` echo is terminal except for [Node.js background retries](/docs/connect/clients/nodejs/#durable-acknowledgement). WebSocket transports only. | | `durable_ack_keepalive_interval_millis` | int (ms) | `200` | Cadence of WebSocket PING the I/O loop sends while there are pending durable confirmations and the producer is idle. `0` or negative disables. | ## Error-handling keys @@ -94,9 +94,9 @@ canonical entries. | `tls_roots` | path | system trust (Node.js: bundled CAs) | Custom CA trust store. | | `tls_roots_password` | string | unset | Trust store password. | | `auto_flush` | bool | `on` | Global on/off for auto-flush triggers. | -| `auto_flush_rows` | int / `off` | `1000` | Row-count flush trigger. | +| `auto_flush_rows` | int / `off` | `1000` | Row-count flush trigger. Node.js: use `0` to disable; `off` is rejected. | | `auto_flush_bytes` | int / `off` | `0` (off) | Byte-size flush trigger. | -| `auto_flush_interval` | int (ms) / `off` | `100` | Time-since-first-row flush trigger (Node.js: since last flush or sender creation). | +| `auto_flush_interval` | int (ms) / `off` | `100` | Time-since-first-row flush trigger (Node.js: since last flush or sender creation). Node.js: use `0` to disable; `off` is rejected. | | `init_buf_size` | size | `64K` | Initial encode buffer capacity; not supported by the Node.js QWP `ws`/`wss` client. | | `max_buf_size` | size | `100M` | Max encode buffer capacity; not supported by the Node.js QWP `ws`/`wss` client. | | `max_name_len` | int | `127` | Local validation cap for table / column names. | diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 27cbc02827..283a21547d 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -102,17 +102,24 @@ GCS, or NFS). - WAL-local durability on the primary is sufficient. - You want minimum steady-state disk usage. - You are running OSS or a build that does not support durable-ack. - (The handshake fails loudly if you opt in but the server cannot - deliver — see below.) + Opting in rejects connection attempts; see [Caveats](#caveats) for the + Node.js background-retry exception. ### Caveats - **Server support is required.** The client sends `X-QWP-Request-Durable-Ack: true` on the upgrade. The server must echo - back `X-QWP-Durable-Ack: enabled`. If it does not — OSS build, - uninitialised primary, missing registry, hitting a replica — the - connect **fails loudly**, by design. Silently waiting for ack frames - that never arrive would let the SF disk fill up. + back `X-QWP-Durable-Ack: enabled`. Without it, for example on an OSS build + or an uninitialised primary, the connection attempt is rejected. This is + normally terminal, subject to the Node.js exception below. +- **Node.js background retries.** Senders with `initial_connect_retry=async` + or `lazy_connect=on` keep retrying unsupported durable acknowledgement + instead of failing initialization. A store-and-forward sender also retries + after its first successful connection. They emit `durable-ack-unavailable` + [connection events](/docs/connect/clients/nodejs/#connection-events). + Monitor these events and buffer usage: continued buffering can fill the + journal or memory queue even though startup succeeded. See + [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). - **Idle keepalive.** The OSS server only flushes pending durable-ack frames during inbound recv events. The client sends a WebSocket PING every `durable_ack_keepalive_interval_millis` (default 200 ms) when From fe424b4ef032daf43050d07823b4579ddaa1be3d Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 00:15:09 +0100 Subject: [PATCH 12/25] docs(nodejs): address QWP client review findings - Store trade IDs as VARCHAR in the store-and-forward and full examples, and document the 2,000,000-value per-sender SYMBOL dictionary cap - Correct rowsAffected, cancel(), max_lifetime_ms, and flush() acceptance semantics, and unwrap QwpReconnectExhaustedError in connectionErrors() - Document the typed reconnect fields, logging hooks, and a table of differences from other clients; merge the sender close semantics into one section - State per-client authentication and durable-ack retry behavior on the shared failover, store-and-forward, and QWP pages, and use one wording for the Node.js memory-mode reconnect exception - Make the store-and-forward error policy table client-agnostic and remove the remaining server-deduplication claims - Fix the UDP client list, separate-connection wording, query overview failover links, quick start client sentence, and changelog entry --- documentation/changelog.mdx | 3 +- .../connect/clients/connect-string.md | 50 +- documentation/connect/clients/nodejs.md | 549 ++++++++++++------ documentation/connect/overview.md | 16 +- .../wire-protocols/qwp-client-behavior.md | 69 ++- .../wire-protocols/qwp-ingress-websocket.md | 21 +- documentation/getting-started/quick-start.mdx | 2 +- .../client-failover/concepts.md | 56 +- .../client-failover/configuration.md | 15 +- .../store-and-forward/concepts.md | 70 ++- .../store-and-forward/configuration.md | 13 +- .../store-and-forward/operating-and-tuning.md | 42 +- .../store-and-forward/when-to-use.md | 30 +- documentation/query/overview.md | 34 +- 14 files changed, 599 insertions(+), 371 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index e42ca7e273..ade7dc3dc9 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -30,7 +30,8 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ. Clarified null handling, ordinary versus background-memory retry policies, HTTP transaction boundaries, replay deduplication, stale-lock recovery, journal capacity allowances, configuration override precedence, and reuse of existing PGWire table schemas. Also clarified cancellation and handshake deadlines, typed pool settings, row validation versus publication failures, acknowledgement error timing, and background durable-ack retries +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ +- [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and [client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 8bf5b5f2c4..ff00905635 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -641,8 +641,9 @@ SF mode and memory-only mode share the same loop. A **running** sender retries a transport outage indefinitely with capped exponential backoff — there is no wall-clock give-up: the whole point of the buffering architecture is that a producer survives an arbitrarily long outage. The -Node.js client is the exception: in memory mode, it gives up after -`reconnect_max_duration_millis`; see below. +exception is a Node.js sender with neither `sf_dir` nor background replay +(`initial_connect_retry=async`, or pooled `lazy_connect=on`), which gives up +after `reconnect_max_duration_millis`; see below. - `reconnect_initial_backoff_millis` — initial wait between reconnect attempts. Backoff grows exponentially up to `reconnect_max_backoff_millis`. @@ -656,9 +657,11 @@ Node.js client is the exception: in memory mode, it gives up after constructor gives up and returns the error. The running loop and the `async` initial connect never consult it. Default: `300000` (5 min). Setting this enables `initial_connect_retry=on` implicitly; see below. - The Node.js client differs: a sender with neither `sf_dir` nor - `initial_connect_retry=async` applies this budget to every outage, and fails - with `QwpReconnectExhaustedError` when it runs out. + The Node.js client differs: a sender with neither `sf_dir` nor background + replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies + this budget to every outage, and fails with `QwpReconnectExhaustedError` + when it runs out. See the + [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - `initial_connect_retry` — whether the client retries the initial connect attempt on failure. - `off` (default, alias `false`) — fail fast on initial connect failure. @@ -666,8 +669,8 @@ Node.js client is the exception: in memory mode, it gives up after thread, up to `reconnect_max_duration_millis`. - `async` — return the `Sender` immediately; the I/O thread retries in the background indefinitely, surfacing terminal failures via the error - inbox. Initial authentication rejection remains terminal; see the Node.js - recovery exception below. + inbox. An authentication rejection before the first successful connection + is terminal; see the authentication note below. **Implicit promotion.** Setting any explicit `reconnect_*` key without also choosing an `initial_connect_retry` mode promotes @@ -685,13 +688,13 @@ Node.js client is the exception: in memory mode, it gives up after This is the shutdown data-loss window. Setting it to `0` skips the drain entirely and drops un-ACKed batches on every clean shutdown. -Authentication rejection (HTTP `401` / `403`) normally stops the reconnect -loop without trying other hosts. The Node.js client makes an exception after -a regular sender's first successful connection: senders with `sf_dir` or -background memory replay (`initial_connect_retry=async` or `lazy_connect=on`) -retry authentication rejections indefinitely. Initial authentication rejection -remains terminal. This exception does not apply to ordinary memory-only senders -or orphan drainers. See +Authentication rejection (HTTP `401` / `403`) never moves the loop to another +host. Before a sender's first successful connection it is terminal in every +client. After that, clients differ: the Java client retries it indefinitely, +the Node.js client does so for senders with `sf_dir` or background replay +(`initial_connect_retry=async`, or pooled `lazy_connect=on`), and the Rust, C, +C++, Python, Go, and .NET clients stop. Query connections and orphan drainers +always stop. See [authentication during failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). ### Egress failover {#egress-failover} @@ -754,7 +757,8 @@ keys so that the Sender and the `QwpQueryClient` can share a single connect string without an "unknown configuration key" error — the Sender does not interpret the values. Range, enum, and type checks happen on the egress side; the Sender silently accepts even a value the -`QwpQueryClient` parser would reject. +`QwpQueryClient` parser would reject. The Node.js `Sender` accepts these keys +too, but logs a warning that it ignores them. - `compression` — result-batch compression the client advertises. Options: `raw` (default — no compression; the client omits the accept-encoding @@ -800,8 +804,9 @@ per-language names. `QuestDBClient.Connect`, `connectQwpNodeClient`).* Every client now leads with a pooled facade, so these keys are a first-contact -concern. The `Sender` and query-client parsers accept and ignore them; the -facade reads them off the string. Each has an equivalent builder setter, and an +concern. The `Sender` and query-client parsers accept and ignore them (the +Node.js `Sender` logs a warning that it ignores them); the facade reads them +off the string. Each has an equivalent builder setter, and an explicit setter always wins over the string. - `sender_pool_min` — senders kept open even when idle. `0` lets the pool close @@ -817,7 +822,8 @@ explicit setter always wins over the string. forever. Default: `60000`. - `max_lifetime_ms` — maximum age of a connection; the housekeeper closes and reopens older ones once idle. `0` means no age limit. Default: `1800000` - (30 min). + (30 min). The Node.js client only closes idle connections above the pool + minimum, and does not recycle the minimum connections. - `housekeeper_interval_ms` — how often the housekeeper checks for idle and over-age connections. Default: `5000`. - `lazy_connect` — when `on`, the pool defers opening its first connection @@ -839,10 +845,10 @@ consumed by the application. :::caution Accepted, but not applied by every client Every client's parser accepts the six `on_*_error` keys below, but only -clients that implement the policy layer act on them. **In the Java reference -client and the Node.js client they are currently accepted no-ops** — setting -`on_write_error=retriable_other` parses cleanly and changes nothing. .NET does -implement them, via `SenderErrorPolicy` and `SenderErrorCategory`. The +clients that implement the policy layer act on them. The Go and .NET clients +apply them. **In the Java reference client and the Node.js, Rust, C, C++, and +Python clients they are currently accepted no-ops**, so setting +`on_write_error=retriable_other` parses cleanly and changes nothing. The category table and precedence model below describe the target contract. ::: diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 28f606bbe6..53963690ef 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -18,20 +18,28 @@ as typed, column-oriented batches. Key capabilities: -- **Ingestion**: a fluent row API and compiled, type-checked object-row - writers, with automatic table creation, schema evolution, batching, and - acknowledgement tracking. -- **Querying**: SQL with typed bind parameters, results streamed as columnar - batches, DDL and DML execution, cancellation, deadlines, and flow control. -- **One pooled client**: `connectQwpNodeClient()` configures ingestion and - queries from one `ws::` connect string, then hands out pooled senders - (`db.borrowSender()`) and query leases (`db.borrowQuery()`). -- **Failover**: multi-host endpoint lists, automatic reconnect, and replay of - unacknowledged rows. -- **Store-and-forward**: a disk journal that keeps accepting rows while - QuestDB is unreachable and survives process restarts. -- **UDP**: fire-and-forget ingestion for metrics where occasional loss is - acceptable. +- **[Ingestion](#data-ingestion)**: a fluent row API and compiled, + type-checked object-row writers, with automatic table creation, schema + evolution, batching, and acknowledgement tracking. +- **[Querying](#querying)**: SQL with typed bind parameters, results streamed + as columnar batches, DDL and DML execution, cancellation, deadlines, and flow + control. +- **[One pooled client](#the-connection-pool)**: `connectQwpNodeClient()` + configures ingestion and queries from one `ws::` connect string, then hands + out pooled senders (`db.borrowSender()`) and query leases + (`db.borrowQuery()`). +- **[Failover](#failover-and-high-availability)**: multi-host endpoint lists, + automatic reconnect, and replay of unacknowledged rows. Replay is at least + once: pair it with table [deduplication](/docs/concepts/deduplication/) for + exactly-once ingestion. +- **[Store-and-forward](#store-and-forward)**: a disk journal that keeps + accepting rows while QuestDB is unreachable and survives process restarts. +- **[UDP](#fire-and-forget-udp)**: fire-and-forget ingestion for metrics where + occasional loss is acceptable. +- **[Error handling](#error-handling)**: typed errors, asynchronous rejection + callbacks, and connection events. The Node.js client differs from the other + QWP clients in a few places; see + [Differences from other clients](#differences-from-other-clients). :::tip Legacy transports @@ -77,8 +85,9 @@ row. Ingestion is asynchronous, so an immediate read may not see it yet. This example assumes `trades` does not exist yet. If it already exists, QWP uses its existing designated timestamp column. For the `trades(ts, ...)` schema in the [PGWire guide](/docs/connect/compatibility/pgwire/nodejs/), -replace `SELECT timestamp` with `SELECT ts` in the query below. The sender's -`at()` calls need no change: they write to the existing designated timestamp. +replace the `timestamp` column with `ts` in every `trades` query on this page. +The sender's `at()` calls need no change: they write to the existing +designated timestamp. ::: @@ -152,9 +161,8 @@ What happens: query handle that is an async iterable of result batches. `batch.rows()` yields one array per row. `query.completion` resolves when the server finishes the query. -4. `db.close()` closes the pools. Idle senders publish any remaining rows and - wait up to five seconds for QuestDB to acknowledge them. `db.close()` - resolves even when that wait times out; see +4. `db.close()` closes the pools, and resolves even if QuestDB has not + acknowledged every row; see [Closing the pooled client](#closing-the-pooled-client). If `trades` did not exist, ingestion creates it automatically with a designated @@ -274,7 +282,7 @@ The `QwpClient` handle has five members: | `borrowQuery()` | `Promise` | Lease an exclusive query connection. Its `close()` returns it to the pool. | | `connect()` | `Promise` | Open the pool minimums. Called for you by `connectQwpNodeClient()`. Safe to retry after a failure. | | `metrics` | `QwpClientMetrics` | Pool counters (`total`, `available`, `leased`, `creating`, `waiting`) for senders and queries. | -| `close()` | `Promise` | Reject new borrows, cancel active queries, close the query connections and idle senders, and wait up to 5 seconds for borrowed senders to be returned. Resolves even if rows are not acknowledged; see [Closing the pooled client](#closing-the-pooled-client). Idempotent. | +| `close()` | `Promise` | Close both pools. Resolves even if rows are not acknowledged; see [Closing the pooled client](#closing-the-pooled-client). Idempotent. | Share one `QwpClient` across your application and close it at shutdown. See [The connection pool](#the-connection-pool) for pool sizing and lease rules. @@ -377,8 +385,11 @@ A QWP connect string has the form `schema::key=value;key=value;`: comma-separated (`addr=a:9000,b:9000`) or by repeating the key. Enclose IPv6 addresses in brackets: `addr=[::1]:9000`. - **Keys** are lowercase and case-sensitive. An unrecognized key fails with - `unknown configuration key: `. Legacy ILP keys such as `retry_timeout` - or `init_buf_size` fail with a hint that names the QWP replacement. + `unknown configuration key: `. Legacy ILP keys fail with a hint: + `retry_timeout` and `tls_ca` name their QWP replacements + (`reconnect_max_duration_millis` and `tls_roots`), and keys with no QWP + equivalent, such as `init_buf_size`, say that they apply only to the legacy + transports. - **Values** end at `;`. Double a semicolon to include it in a value: `password=p;;ssw;;rd` sets the password to `p;ssw;rd`. The trailing `;` is optional. @@ -437,14 +448,30 @@ where `qwp` has the sections `webSocket`, `session` (the equivalent of :::caution A typed `reconnect` object replaces the connect-string keys `ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces -the whole reconnect policy parsed from `reconnect_*` keys, and +the whole reconnect policy parsed from the `reconnect_*`, +`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and `egressSession.reconnect` replaces the policy parsed from `failover*` keys. Fields you leave out of the object take the built-in defaults, not the values from the connect string. When you supply the object, for example to register -`onEvent`, set every bound you rely on in it, such as `maxDurationMs`. +`onEvent`, set every bound you rely on in it. ::: +The object's fields and the connect-string keys they replace: + +| Field | Ingestion key, default | Query key, default | +|---|---|---| +| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8` | +| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | +| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | +| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | +| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | +| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | +| `onEvent` | None | None | + +`egressSession: { reconnect: false }` is the typed equivalent of +`failover=off`. + ## Authentication and TLS @@ -521,42 +548,23 @@ The pooled client reports connection setup failures as the `cause` of a | Path | Status | Workaround | |---|---|---| | OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider and has no callback to refresh a token. | Obtain an access token from your identity provider, pass it as `token=...`, and create a new client before the token expires. See [OpenID Connect](/docs/security/oidc/). | -| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. | Close the client and create a new one with the new token. | +| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and memory-mode senders then fail, while senders with `sf_dir` or background replay keep retrying and buffering (see [Connection-level errors](#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | | Mutual TLS (client certificates) | Not supported. QuestDB does not negotiate client certificates. | Use token or basic authentication over `wss`. | | ILP JWK authentication | Not available for QWP. `auth`, `jwk`, `token_x`, and `token_y` are rejected on `ws`/`wss`. | Use token or basic authentication. | ### Production example: TLS, token, and multiple hosts -A typical Enterprise deployment combines `wss`, a token, and several hosts: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const token = process.env.QDB_TOKEN; -if (!token) throw new Error("QDB_TOKEN is not set"); +A typical Enterprise deployment combines `wss`, a token, and several hosts in +one connect string: -const db = await connectQwpNodeClient( - "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + - `token=${token};` + - "tls_roots=/etc/ssl/questdb-ca.pem;" + - // Start, and ingest, even while no replica is reachable. - "query_pool_min=0;", - { - // Queries run on replicas only, with no fallback to the primary (see - // "Multiple endpoints"). Set target here: in the connect string it also - // applies to ingestion. - egress: { target: "replica" }, - }, -); -try { - // borrow senders and query leases -} finally { - await db.close(); -} +```text +wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN; ``` -With `query_pool_min=0`, the client starts while no replica is reachable, and -a query borrowed during that time rejects with `QwpPoolResourceError`. +Add `tls_roots=/path/to/ca.pem;` when the servers use a private CA. See +[Multiple endpoints](#multiple-endpoints) for routing queries to replicas, and +the [full example](#full-example-ingestion-and-querying-with-failover) for a +complete program with this configuration. ## The connection pool @@ -601,24 +609,10 @@ A long-running producer can keep its borrow for its whole lifetime and call `flush()` between batches. Size `sender_pool_max` to the number of producers that hold a sender at the same time. -:::note Pooled sender close semantics - -`close()` on a borrowed sender flushes completed rows, discards an unfinished -row with a warning, and returns the sender to the pool. It does not close the -WebSocket or wait for acknowledgements by default. With `awaitServerAck: true` -or `awaitDurableAck: true`, the flush performed by `close()` waits for its -acknowledgement too. To confirm delivery of all previously published rows -before returning the sender, call `flush()` and then -`waitForAcknowledged(sender.publishedSequence)`; see -[Awaiting acknowledgements](#awaiting-acknowledgements). - -When a borrowed sender's `close()` fails, the pool discards that sender and -opens a new one for the next borrow. Because QuestDB reports rejected batches -asynchronously, a sender can fail after its `close()` already succeeded: the -error then surfaces on the `flush()` or `close()` of the next borrower, and the -pool replaces the sender after that. See [Ingestion errors](#ingestion-errors). - -::: +`close()` on a borrowed sender flushes its completed rows and returns it to +the pool without waiting for acknowledgements. See +[Closing a sender](#closing-a-sender) for how to wait for them, and for what +happens when a close fails. ### Borrowing a query lease @@ -672,7 +666,7 @@ Starting a second query on a lease while one is still active throws | `query_pool_max` | `4` | Maximum query connections, which also caps concurrent queries. | | `acquire_timeout_ms` | `5000` | How long a borrow waits when the pool is at its maximum, before rejecting with `QwpPoolAcquireTimeoutError`. | | `idle_timeout_ms` | `60000` | Idle time before an excess connection is closed. `0` keeps idle connections. | -| `max_lifetime_ms` | `1800000` | Age at which an idle connection is recycled. `0` disables recycling. | +| `max_lifetime_ms` | `1800000` | Age at which an idle connection above the pool minimum is closed. Connections kept open by `sender_pool_min` and `query_pool_min` are never recycled, so this does not rotate a pool that is at its minimum. `0` disables it. | | `housekeeper_interval_ms` | `5000` | How often the housekeeper checks for idle and over-age connections. Minimum `100`. | | `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | | `lazy_connect` | `off` | Start without connecting. See below. | @@ -750,12 +744,13 @@ the client closes before QuestDB becomes reachable; see senders to be returned. A sender still borrowed after that stays open, and its owner must `close()` it. -`db.close()` resolves even when an acknowledgement does not arrive in time. It -reports the timeout to `ingressSession.onError` as a non-terminal -`QwpIngressAckTimeoutError`, which is logged as a warning by default. Without -`sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they stay in -the journal, and the next sender on the same directory replays them. To know -that QuestDB accepted every row before shutting down, wait for the +`db.close()` resolves even when an acknowledgement does not arrive in time. +Without `sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they +stay in the journal, and the next sender on the same directory replays them. +The client usually reports the timeout to `ingressSession.onError` as a +non-terminal `QwpIngressAckTimeoutError`, logged as a warning by default, but +the report is best-effort: do not rely on it to detect unacknowledged rows. To +know that QuestDB accepted every row before shutting down, wait for the acknowledgement before returning each sender (see [Awaiting acknowledgements](#awaiting-acknowledgements)), or use [store-and-forward](#store-and-forward). @@ -778,7 +773,7 @@ state, so borrow one sender per producer (see [Concurrency](#concurrency)). [Null values](#null-values) for non-nullable defaults). 4. Close the row with `at(timestamp, unit)` or `atNow()`, and `await` the returned promise. It rejects if an auto-flush triggered by the row fails. -5. Repeat from step 2, and call `flush()` to send staged rows. +5. Repeat from step 2, and call `flush()` to publish staged rows. 6. `close()` the sender when done. ```typescript @@ -804,6 +799,16 @@ try { } ``` +`flush()` resolving means the rows were published, not that QuestDB accepted +them. QuestDB reports a rejected batch, such as one with a value of the wrong +type for an existing column, after `flush()` has resolved: to the +`onSenderError` callback, which only logs it by default, and as a rejection of +`waitForAcknowledged()`. When your code must know that QuestDB accepted the +rows, wait for the acknowledgement after flushing, with +`await sender.waitForAcknowledged(sender.publishedSequence)`. See +[Awaiting acknowledgements](#awaiting-acknowledgements) and +[Ingestion errors](#ingestion-errors). + Tables and columns are created automatically, with the column types listed below. Table and column names are validated locally with QuestDB's rules (at most 127 UTF-8 bytes by default, see `max_name_len`), and column names are @@ -867,6 +872,19 @@ Names that differ from what you might expect: - `geohashColumn()` takes raw bits only. Base-32 geohash text is accepted by a compiled writer's `geohash()` field. +:::caution SYMBOL is for bounded sets of values + +Use SYMBOL for values from a bounded set, such as tickers, sides, or venues. +Each sender keeps every distinct SYMBOL value it has sent, across all tables +and columns, in a dictionary that holds at most 2,000,000 values and is not +cleared while the sender lives. Once it is full, every later `at()`, `flush()`, +and `close()` on that sender fails with an `Error`, and the sender's staged +rows are lost. Store unique or high-cardinality values, such as trade or order +IDs, as VARCHAR with `stringColumn()`, or as UUID with `uuidColumn()`. See +[Symbol](/docs/concepts/symbol/). + +::: + The standalone `Sender` class exposes only `symbol`, `stringColumn`, `booleanColumn`, `floatColumn`, `intColumn`, `timestampColumn`, `arrayColumn`, `decimalColumn`, and `decimalColumnText`. Its `writer()` method supports every @@ -986,7 +1004,9 @@ try { ``` `at(value, unit)` accepts an integer `number` or a `bigint` with unit `"us"` -(the default), `"ms"`, or `"ns"`. Nanoseconds require a `bigint`, because epoch +(the default), `"ms"`, or `"ns"`. `Date.now()` returns milliseconds, so always +pass `"ms"` with it: without a unit, the value is read as microseconds and the +row lands in January 1970. Nanoseconds require a `bigint`, because epoch nanoseconds exceed the safe integer range. When the table does not exist yet, `"ns"` creates a `TIMESTAMP_NS` designated timestamp and the other units create a microsecond `TIMESTAMP`. An auto-created designated timestamp column is named @@ -1183,7 +1203,7 @@ Schema fields: | `float32()` | FLOAT | `number` | | `float64()`, `double()` | DOUBLE | `number` | | `timestamp(unit)` | TIMESTAMP or TIMESTAMP_NS | `number` or `bigint`; `"ns"` requires `bigint` | -| `designatedTimestamp(unit)` | designated TIMESTAMP | As `timestamp(unit)`, required in every row. At most one per schema | +| `designatedTimestamp(unit)` | Designated TIMESTAMP, or TIMESTAMP_NS with `"ns"` | As `timestamp(unit)`, required in every row. At most one per schema. The field's key names the row property only: the value always goes to the table's designated timestamp, which is named `timestamp` when QWP creates the table | | `date()` | DATE | Epoch milliseconds | | `binary()` | BINARY | `Uint8Array` | | `uuid()` | UUID | Canonical UUID `string`, 16 big-endian bytes, or `{ low, high }` | @@ -1236,32 +1256,37 @@ then rejects with `QwpMemoryReplayAppendTimeoutError`. Tune the cap with `sf_max_total_bytes` and the wait with `sf_append_deadline_millis`; without `sf_dir` they size the memory queue. Watch `sender.metrics.ingress` (`memoryReplayUsedBytes`, `totalMemoryReplayBackpressureStalls`) to detect -backpressure before it blocks. +backpressure before it blocks. `metrics` is available on pooled senders and on +senders from `connectQwpNodeSender()`, not on the standalone `Sender` class. **Oversized rows.** When the sender connects, QuestDB advertises the largest batch it accepts: about 2 MiB on a default server, set by -`http.recv.buffer.size`. With `sf_dir`, a batch must also fit in -`sf_max_segment_bytes`. A row too large to fit in one batch fails the +`http.recv.buffer.size`. A batch must also fit in `sf_max_segment_bytes`, +which defaults to 4 MiB with `sf_dir` and applies without `sf_dir` only when +you set it. A row too large to fit in one batch fails the `flush()`, or the `at()` whose auto-flush sends it, with `QwpBatchTooLargeError` before anything is sent. The staged rows are kept, so every later flush fails the same way, and `close()` discards them and rejects with the same error. Call `reset()` to drop every row staged since the last flush, then write the other rows again. -**Closing.** `close()` on a standalone sender publishes completed rows and waits -up to `close_flush_timeout_millis` (5 seconds by default) for their -acknowledgement. `0` or a negative value skips the wait. An unfinished row is -discarded with a warning. On a borrowed sender, `close()` flushes and returns -the sender to the pool without waiting for acknowledgements by default; see -[Borrowing a sender](#borrowing-a-sender). - -If the acknowledgement does not arrive in time, `close()` on a standalone -sender rejects with `QwpSenderCloseTimeoutError`. Its `targetSequence` is the -last published sequence and its `acknowledgedSequence` is how far QuestDB -acknowledged. Without `sf_dir`, the unacknowledged rows may be lost. With -`sf_dir`, they stay in the journal for the next sender on that directory. A -rejection in `finally` replaces any error the `try` block threw, so catch it -there when that matters: +### Closing a sender + +`close()` publishes the sender's completed rows and discards an unfinished row +with a warning. The rest depends on how you created the sender. With +transactions on, see also [Transactions](#transactions). + +#### Standalone sender + +`close()` waits up to `close_flush_timeout_millis` (5 seconds by default) for +QuestDB to acknowledge every published row, then closes the connection. `0` or +a negative value skips the wait. If the acknowledgement does not arrive in +time, `close()` rejects with `QwpSenderCloseTimeoutError`. Its +`targetSequence` is the last sequence `close()` waited for, and its +`acknowledgedSequence` is how far QuestDB acknowledged. Without `sf_dir`, the +unacknowledged rows may be lost. With `sf_dir`, they stay in the journal for +the next sender on that directory. A rejection in `finally` replaces any error +the `try` block threw, so catch it there when that matters: ```typescript import { QwpSenderCloseTimeoutError, Sender } from "@questdb/nodejs-client"; @@ -1289,6 +1314,23 @@ try { } ``` +#### Borrowed sender + +`close()` returns the sender to the pool without closing its connection and, +by default, without waiting for acknowledgements. With `awaitServerAck: true` +or `awaitDurableAck: true`, the flush performed by `close()` waits for its +acknowledgement too. To confirm delivery of every row the sender published +before returning it, call `flush()` and then +`waitForAcknowledged(sender.publishedSequence)`; see +[Awaiting acknowledgements](#awaiting-acknowledgements). + +When a borrowed sender's `close()` fails, the pool discards the sender and +opens a new one for the next borrow. Because QuestDB reports rejected batches +asynchronously, a sender can fail after its `close()` already succeeded: the +error then surfaces on the next borrower's auto-flushing `at()`, `flush()`, or +`close()`, and the pool replaces the sender after that. See +[Ingestion errors](#ingestion-errors). + ### Awaiting acknowledgements QuestDB acknowledges ingested batches asynchronously. Every published frame gets @@ -1436,11 +1478,13 @@ order, and acknowledged segments are deleted. Before ingesting, create a deduplicated table while QuestDB is reachable. Use both the event timestamp and a stable, source-assigned trade ID as upsert keys: distinct trades can share a millisecond timestamp, symbol, and side. +Store the trade ID as VARCHAR, not SYMBOL: every trade has its own ID, and +SYMBOL is for [bounded sets of values](#column-methods). ```questdb-sql CREATE TABLE trades_sf ( timestamp TIMESTAMP, - trade_id SYMBOL, + trade_id VARCHAR, symbol SYMBOL, side SYMBOL, price DOUBLE, @@ -1467,7 +1511,7 @@ try { try { await sender .table("trades_sf") - .symbol("trade_id", event.tradeId) + .stringColumn("trade_id", event.tradeId) .symbol("symbol", "ETH-USD") .symbol("side", "buy") .doubleColumn("price", 2615.54) @@ -1494,12 +1538,18 @@ over the directory's lock (see Lock recovery below). `/-0`, `/-1`, and so on. `sender_id` defaults to `default` and may contain letters, digits, `_`, and `-`. Give every process its own `sender_id`; a second live process on the same - directory fails with `QwpReplayStoreLockedError`. -- **Durability.** `sf_durability=memory` (the connect-string default) relies on - the operating system to write the journal, which survives a process crash but - not a power loss. `periodic` checkpoints in the background every - `sf_sync_interval_millis` (5 seconds). `append` makes every append durable - before `flush()` resolves. + directory fails with `QwpReplayStoreLockedError`. A pooled client also + drains any of its own `-` journals that no pooled sender + holds, such as those left by a larger pool before a restart, without + `drain_orphans`. +- **Durability.** `sf_durability` sets how the journal reaches the disk. It is + unrelated to memory mode, which means running without `sf_dir`. `memory` + (the connect-string default) relies on the operating system to write the + journal, which survives a process crash but not a power loss. `periodic` + checkpoints in the background every `sf_sync_interval_millis` (5 seconds). + `append` makes every append durable before `flush()` resolves, which adds a + disk sync to every flush: on a producer that flushes often, prefer `periodic` + or larger batches. - **Capacity.** With `sf_dir`, `sf_max_total_bytes` (10 GiB by default) is a journal size target, not a hard disk limit. Transaction-closing batches can reserve extra segments so a full journal does not block the commit needed @@ -1539,8 +1589,8 @@ over the directory's lock (see Lock recovery below). failed pooled sender, sends it again and fails the same way. See [Ingestion errors](#ingestion-errors) for recovery. - **Orphans.** With `drain_orphans=on`, a sender also adopts and drains - journals left under the same `sf_dir` by processes that crashed, up to - `max_background_drainers` (4) at a time. + journals with other `sender_id` values left under the same `sf_dir` by + processes that crashed, up to `max_background_drainers` (4) at a time. A frame appended to the journal but not acknowledged before a crash is sent again, so delivery is at least once: @@ -1706,9 +1756,14 @@ A `QwpResultBatch` has: code (compare it with the exported `QWP_COLUMN_TYPE` constants). - `rows()`: a generator that yields one array of values per row. - `get(rowIndex, columnIndex)`: one value. +- `batchSequence`: the batch's position in the result, starting at `0n`. Batch objects stay valid after iteration moves on, so you can keep them. +With the default `failover=on`, a lost connection can make the client run the +query again from its first batch. If your loop accumulates rows, reset them +when `batch.batchSequence === 0n`; see [Query failover](#query-failover). + ### Reading result values Values arrive as these JavaScript types: @@ -1727,7 +1782,7 @@ Values arrive as these JavaScript types: | BINARY | `Uint8Array` | | IPv4 | `number`, as a signed 32-bit integer: `192.168.0.1` arrives as `-1062731775`. Use `value >>> 0` for the unsigned address | | UUID | `{ low: bigint, high: bigint }`, the unsigned low and high 64-bit halves | -| LONG256 | `{ words: [bigint, bigint, bigint, bigint] }`, least significant word first | +| LONG256 | `{ words: [bigint, bigint, bigint, bigint] }`, least significant word first. Each word is a signed 64-bit value: `BigInt.asUintN(64, word)` gives its unsigned value | | GEOHASH | `{ bits: bigint, precisionBits: number }` | | DECIMAL with a precision of 10 or more | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | | DOUBLE[], DOUBLE[][], ... | `{ dimensions: number[], values: number[] }` with values in row-major order | @@ -1750,7 +1805,9 @@ Convert the column in SQL instead: - An untyped `NULL` literal, as in `SELECT NULL`: give it a type, for example `NULL::double`. -Converting common types: +`JSON.stringify()` throws on `bigint`, which LONG, TIMESTAMP, DATE, UUID, +LONG256, and DECIMAL values contain, so convert rows before serializing them, +as `toJson()` does below. Converting common types: ```typescript // TIMESTAMP (bigint microseconds) to Date. Drops sub-millisecond precision. @@ -1773,9 +1830,14 @@ function uuidToString({ low, high }: { low: bigint; high: bigint }): string { const ipv4ToString = (value: number) => [24, 16, 8, 0].map((shift) => ((value >>> 0) >>> shift) & 0xff).join("."); +// Any row or value to JSON, with bigint values as decimal strings +const toJson = (value: unknown) => + JSON.stringify(value, (_key, v) => (typeof v === "bigint" ? v.toString() : v)); + console.log(toDate(1723000000000000n).toISOString()); console.log(uuidToString({ low: 13485158461794337056n, high: 11465204444048149893n })); console.log(ipv4ToString(-1062731775)); +console.log(toJson([1723000000000000n, "ETH-USD", 2615.54])); ``` ### Bind parameters @@ -1842,6 +1904,12 @@ There is no setter for BINARY, IPv4, or arrays. Bind IPv4 as a string and cast it in SQL (`WHERE ip = $1::ipv4` with `setVarchar`), and pass array values as SQL literals. +Decimal and geohash setters take the scale or precision before the value, the +reverse of the matching column methods: `setDecimal64(index, scale, unscaled)` +but `decimal64Column(name, unscaled, scale)`, and +`setGeohash(index, precisionBits, value)` but +`geohashColumn(name, bits, precisionBits)`. + ### DDL and DML statements `CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` go through the @@ -1855,8 +1923,10 @@ including DDL and DML. QuestDB may have applied an `INSERT` before its `exec-done` response was lost, so replay can insert it again. A transport error does not prove the statement failed. Use a separate client with `failover=off` for non-idempotent statements, as below, and verify an -uncertain outcome before retrying manually. Alternatively, make the SQL -idempotent; see [Query failover](#query-failover). +uncertain outcome before retrying manually. Alternatively, make the statement +idempotent so that a second run changes nothing, for example +`CREATE TABLE IF NOT EXISTS`, or an `INSERT` of rows with stable key values +into a table with [deduplication](/docs/concepts/deduplication/). ::: @@ -1880,8 +1950,9 @@ try { for (const sql of statements) { const statement = await lease.query(sql); const completion = await statement.completion; - if (completion.kind === "exec-done") { - console.log(`${sql.slice(0, 20)}...: ${completion.rowsAffected} rows`); + // rowsAffected counts rows only for INSERT; see the table below. + if (completion.kind === "exec-done" && sql.startsWith("INSERT")) { + console.log(`inserted ${completion.rowsAffected} rows`); } } } catch (error) { @@ -1898,7 +1969,13 @@ try { | `completion.kind` | Returned for | Fields | |---|---|---| | `"result-end"` | Queries that return rows | `totalRows` (`bigint`) | -| `"exec-done"` | DDL and DML | `rowsAffected` (`bigint`, `0` for DDL), `operationType` (QuestDB's numeric statement type) | +| `"exec-done"` | DDL and DML | `rowsAffected` (`bigint`, rows written by an `INSERT`), `operationType` (QuestDB's numeric statement type) | + +Only `INSERT` reports a row count in `rowsAffected`. For an `UPDATE` on a WAL +table, the default, it currently holds a transaction number rather than the +number of rows changed, so do not use it to check whether an `UPDATE` matched +any rows. DDL reports `0`, except `TRUNCATE`, which currently reports +`18446744073709551615n`. Statements run in order on one lease, because each is awaited before the next starts, so a `CREATE TABLE` is complete before the `INSERT` that follows it. @@ -1909,35 +1986,34 @@ A query ends early in four ways: - **Deadline.** Set a default with `egressSession: { queryTimeoutMs }`, or per query with `timeoutMs`. On expiry, iteration and `completion` reject with - `QwpEgressQueryTimeoutError` and the client sends a cancel to QuestDB. -- **Cancel.** `await query.cancel()` sends a request to stop; it does not wait - for QuestDB to stop. When QuestDB processes the cancel, iteration and - `completion` reject with `QwpEgressQueryError` whose `status` is `0x0a` - (CANCELLED). A query that finishes before the cancel is processed, or a DDL - or DML statement, completes normally instead. + `QwpEgressQueryTimeoutError`, and the client cancels the query in the + background. - **Leaving the loop.** `break`, `return`, or an exception inside `for await` - cancels the query, and `completion` rejects with - `QwpEgressQueryAbandonedError`. + cancels the query in the background, and `completion` rejects with + `QwpEgressQueryAbandonedError`. Use it to stop reading a result at once. +- **Cancel.** `await query.cancel()` asks QuestDB to stop and returns without + waiting for it. Keep consuming the result afterwards: QuestDB acts on the + cancel only while the result is moving, and iteration and `completion` then + reject with `QwpEgressQueryError` whose `status` is `0x0a` + (`QWP_STATUS.CANCELLED`). If you stop consuming and only await `completion`, + it rejects with `QwpEgressQueryCancelTimeoutError` after + `query_close_timeout_ms` (5 seconds), and the client closes the connection. + A query that finishes before QuestDB processes the cancel, or a DDL or DML + statement, completes normally instead. - **Waiting without cancelling.** `await query.awaitCompletion(timeoutMs)` resolves `false` when the wait times out and leaves the query running. `query.isDone()` reports whether the query has ended. -QuestDB acts on a cancel between result batches, but while a query streams -without a [credit window](#flow-control), it may not read the cancel until the -whole result is sent. Set `initial_credit` in the connect string or -`initialCredit` per query to improve cancellation responsiveness while -streaming. For example, start with 1 MiB; the client replenishes it as your -loop consumes batches. Credit does not interrupt expensive work before the -next batch or bound cancellation latency. - -With or without a credit window: - -- After an explicit cancel, iteration and `completion` can reject with - `QwpEgressQueryCancelTimeoutError` instead of the `0x0a` rejection: the - client waits `query_close_timeout_ms` (5 seconds) for QuestDB to stop, then - closes the connection. -- After a deadline or an early exit from the loop, returning the lease can take - up to twice `query_close_timeout_ms`. +QuestDB checks for a cancel between result batches, while it is sending. Set a +[credit window](#flow-control), with `initial_credit` in the connect string or +`initialCredit` per query, so that QuestDB pauses when your loop falls behind +and stops at the next batch after a cancel, a deadline, or an early exit. +Start with 1 MiB; the client replenishes it as your loop consumes batches. +Without a credit window, QuestDB streams as fast as the network allows: it may +send the whole result before it reads a cancel, so a cancelled query can +complete normally, and returning the lease after an early exit waits while the +rest of the result streams. A credit window does not interrupt expensive work +before the next batch is ready. ```typescript import { @@ -1951,7 +2027,7 @@ try { try { const query = await lease.query( "SELECT symbol, avg(price) FROM trades SAMPLE BY 1m", - // Credit limits streaming ahead, not cancellation latency. + // With credit, QuestDB stops at the next batch after the deadline. { timeoutMs: 5_000, initialCredit: 1024 * 1024 }, ); for await (const batch of query) { @@ -1975,7 +2051,7 @@ cancellation, and another `query()` on the same lease throws `close()` and borrow a new one for the next query. `close()` waits up to `query_close_timeout_ms` (5 seconds) for QuestDB to confirm the cancellation. If it does not, `close()` discards the connection, which can take as long -again. +again, so returning the lease can take up to twice `query_close_timeout_ms`. ### Flow control @@ -2009,8 +2085,9 @@ try { With `autoCredit: false`, call `query.grantCredit(bytes)` yourself. To cap the rows in each batch, set `max_batch_rows` (1 to 1,048,576). A credit window -also lets QuestDB act on a cancel or a deadline promptly; see -[Cancellation and timeouts](#cancellation-and-timeouts). +also lets QuestDB stop at the next batch after a cancel, a deadline, or an +early exit; see [Cancellation and timeouts](#cancellation-and-timeouts) for +what your loop must do after `cancel()`. ### Zero-copy result views @@ -2147,7 +2224,7 @@ When `onSenderError` is not set, rejections are logged: retriable ones at client's protocol handling, and an exception thrown by a callback is contained. A standalone sender's `close()` can also reject, with `QwpSenderCloseTimeoutError`, when its rows are not acknowledged in time; see -[Flushing](#flushing). +[Closing a sender](#closing-a-sender). `QwpSenderError` fields: @@ -2296,19 +2373,25 @@ The pooled client wraps every failure to open a connection, from `connectQwpNodeClient()`, `db.connect()`, `borrowSender()`, or `borrowQuery()`, in a `QwpPoolResourceError`. Unwrap its `cause` before checking for a specific error. When `addr` lists several hosts, the cause is a -`QwpFailoverError` whose `attempts` hold the error of each endpoint: +`QwpFailoverError` whose `attempts` hold the error of each endpoint. When +initial-connect retry is on, for example with a `failover_*` or `reconnect_*` +key or a typed `reconnect` object, the cause is a `QwpReconnectExhaustedError` +instead, and its own `cause` holds the last attempt's error: ```typescript import { connectQwpNodeClient, QwpFailoverError, QwpPoolResourceError, + QwpReconnectExhaustedError, QwpUpgradeError, } from "@questdb/nodejs-client"; // The errors behind a failed connection, one per endpoint tried. function connectionErrors(error: unknown): unknown[] { - const cause = error instanceof QwpPoolResourceError ? error.cause : error; + let cause = error instanceof QwpPoolResourceError ? error.cause : error; + // With initial-connect retry on, the last attempt's error is wrapped. + if (cause instanceof QwpReconnectExhaustedError) cause = cause.cause; return cause instanceof QwpFailoverError ? cause.attempts.map((attempt) => attempt.error) : [cause]; @@ -2336,12 +2419,27 @@ walk because credentials are assumed to be shared across the cluster. After a successful connection, regular senders with `sf_dir` or background memory replay (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying authentication rejections indefinitely. This lets buffered data drain once -server-side authentication is restored. Memory-only senders without those +server-side authentication is restored. Memory-mode senders without those settings, and orphan drainers, do not have this exception. See -[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) +for how other clients behave. Endpoints in error messages have any embedded credentials removed. +### Logging + +The client writes its own messages to the console by default, at the `error`, +`warn`, and `info` levels. To route a sender's messages, such as warnings about +rows discarded on close, through your logger, pass a `(level, message)` +function: `{ sender: { log } }` as the second argument of +`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. The function +also receives `debug` messages, one per staged row, so filter by level. + +Rejected batches and session errors go to `ingressSession.onSenderError` and +`ingressSession.onError`. Their defaults log to the console, so replace both to +route them through your logger. Some messages from other parts of the client, +such as store-and-forward recovery, always go to the console. + ## Failover and high availability :::note Enterprise @@ -2367,14 +2465,33 @@ queries. Ingestion always needs the primary: replicas refuse writes, and the sender walks the list until it finds the current primary. Queries can use any node. `target` selects which roles queries accept: `any` (the default), `primary`, or -`replica`. It is a strict filter, not a preference: with `replica`, queries -never fall back to the primary, and they fail when no replica is reachable, -including against a single open source server. Because the pooled client opens -a query connection at startup, `connectQwpNodeClient()` then rejects too, with -a `QwpPoolResourceError` whose `cause` is a `QwpRoleMismatchError`, or a +`replica`. Set it with the typed `egress` option, as below: in the connect +string, `target` also filters ingestion (see the caution that follows). It is a +strict filter, not a preference: with `replica`, queries never fall back to the +primary, and they fail when no replica is reachable, including against a single +open source server. Because the pooled client opens a query connection at +startup, `connectQwpNodeClient()` then rejects too, with a +`QwpPoolResourceError` whose `cause` is a `QwpRoleMismatchError`, or a `QwpFailoverError` holding one per endpoint when `addr` lists several hosts. -Set `query_pool_min=0` to start without a replica. `zone` prefers endpoints in -the same zone. +With query failover retrying the initial connection, that error is wrapped in a +`QwpReconnectExhaustedError`. To start without a replica, also set +`query_pool_min=0`; queries borrowed before a replica is reachable then reject +with `QwpPoolResourceError`: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// Queries run on replicas only. Ingestion still follows the primary. +const db = await connectQwpNodeClient( + "wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN;" + + // Start, and ingest, even while no replica is reachable. + "query_pool_min=0;", + { egress: { target: "replica" } }, +); +await db.close(); +``` + +`zone` prefers endpoints in the same zone. :::caution `target` in the connect string also filters ingestion @@ -2396,20 +2513,27 @@ jitter, then resends every unacknowledged batch: |---|---|---| | `reconnect_initial_backoff_millis` | `100` | First retry delay. | | `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | -| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in memory mode. `0` removes the limit. | +| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in memory mode, without `sf_dir` or background replay. `0` removes the limit. | | `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | Whether the sender gives up depends on the mode (see [Flushing](#flushing)): -- **Memory mode** retries for up to `reconnect_max_duration_millis` per outage. - When the budget runs out, the sender fails permanently with - `QwpReconnectExhaustedError`, and its unsent rows are lost. -- **Background memory mode** (`initial_connect_retry=async`) and - **store-and-forward** (`sf_dir`) retry indefinitely. - -Setting any `reconnect_*` key also makes the first connection retry within the -budget, as if `initial_connect_retry=on`. Set `initial_connect_retry=off` -explicitly to keep a fail-fast start. +- **Memory mode**, without `sf_dir` or background replay, retries for up to + `reconnect_max_duration_millis` per outage. When the budget runs out, the + sender fails permanently with `QwpReconnectExhaustedError`, and its unsent + rows are lost. The Java reference client retries indefinitely in this mode + instead. +- **Background memory mode** (`initial_connect_retry=async`, or + `lazy_connect=on` on the pooled client) and **store-and-forward** + (`sf_dir`) retry indefinitely. + +Setting any `reconnect_*` key also makes a sender's first connection retry +within the budget, as if `initial_connect_retry=on`. Set +`initial_connect_retry=off` explicitly to keep a fail-fast start. The keys do +not apply to query connections: the pooled client still opens its query pool +at startup, so `connectQwpNodeClient()` fails fast while QuestDB is down unless +you also set `query_pool_min=0` or enable query retries (see +[Connection events](#connection-events)). Replay after a reconnect is at least once: a batch that QuestDB committed just before the connection dropped is sent again. @@ -2549,7 +2673,9 @@ disables the reconnect wrapper. As with other session options, an explicit | `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | `reconnected` and `failed-over` are mutually exclusive: code that tracks the -current node must handle both. +current node must handle both. A query that reconnects runs again from its +first batch; see [Query failover](#query-failover) for resetting accumulated +rows. No event marks a terminal failure. When a sender stops retrying, because its reconnect budget ran out or the error cannot be retried, @@ -2604,12 +2730,13 @@ every key. The Node.js client's defaults and deviations: | `request_durable_ack` | `off` | Enterprise. | | `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | | `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | -| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in memory mode. `0` removes it. | +| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in memory mode, without `sf_dir` or background replay. `0` removes it. | +| `max_frame_rejections`, `poison_min_escalation_window_millis` | `4`, `300000` | Poison-frame detector: rejections of one batch, and the minimum time they must span, before the sender stops. | | `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | | `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | | `sf_durability` | `memory` | `memory`, `periodic`, or `append`. | | `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](#store-and-forward), not a hard disk limit; memory queue cap without `sf_dir`. | -| `sf_max_segment_bytes` | `4m` | Journal segment size, which also caps a batch. | +| `sf_max_segment_bytes` | `4m` with `sf_dir`, none without | Journal segment size, which also caps a batch. | | `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | | `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. | | `target`, `zone` | `any`, none | Endpoint role and zone preference. Apply to ingestion too. | @@ -2617,6 +2744,7 @@ every key. The Node.js client's defaults and deviations: | `compression`, `compression_level` | `raw`, `1` | Query result compression. | | `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | | `client_id` | `typescript/` | Sent to the server for diagnostics. | +| `error_inbox_capacity`, `connection_listener_inbox_capacity` | `256`, `64` | Queues for rejection callbacks and connection events. | | Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client only. | The @@ -2625,6 +2753,26 @@ covers every type and option. The [QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) in the client repository describes the delivery semantics in depth. +### Differences from other clients + +The Node.js client differs from the Java reference client, and from the shared +[connect string reference](/docs/connect/clients/connect-string/), in these +places: + +| Area | Node.js behavior | +|---|---| +| Outage budget | A sender without `sf_dir` or background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. See [Ingestion reconnect](#ingestion-reconnect). | +| `target` and `zone` | Also apply to ingestion. Set a query-only role with the typed `egress.target` option. See [Multiple endpoints](#multiple-endpoints). | +| Authentication rejected after a first connection | Senders with `sf_dir` or background replay keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | +| Durable acknowledgement unavailable | Background senders keep retrying from startup, and store-and-forward senders after their first connection, emitting `durable-ack-unavailable`. See [Durable acknowledgement](#durable-acknowledgement). | +| `sf_durability` | Also accepts `append`. | +| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Store-and-forward](#store-and-forward). | +| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Store-and-forward](#store-and-forward). | +| `max_lifetime_ms` | Closes idle connections above the pool minimum only. Connections at the minimum are not recycled. | +| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush. Size values take single-letter suffixes only. `tls_roots` must be PEM. `init_buf_size` and `max_buf_size` are rejected. | +| Defaults | `connect_timeout` is `15000`, `poison_min_escalation_window_millis` is `300000`, and `close_flush_timeout_millis` is `5000`. | +| `on_*_error` keys | Accepted but not applied. | + ## Migration ### From ILP to QWP @@ -2651,8 +2799,11 @@ connect string and calling `connect()`: | Column types | ILP types | More types, subject to [column-method](#column-methods) and [array](#arrays) support | Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, -`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss`, -with a hint naming the replacement. To keep ILP-sized batches, set +`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss` +with a hint. It names the replacement where there is one +(`retry_timeout` becomes `reconnect_max_duration_millis`, and `tls_ca` becomes +`tls_roots`), and otherwise says that the key applies only to ILP or that QWP +negotiates the setting itself. To keep ILP-sized batches, set `auto_flush_rows` and `auto_flush_interval` explicitly. Migrate one sender at a time: ILP and QWP senders can run side by side. @@ -2731,7 +2882,7 @@ from [Store-and-forward](#store-and-forward)): ```questdb-sql CREATE TABLE IF NOT EXISTS trades_sf ( timestamp TIMESTAMP, - trade_id SYMBOL, + trade_id VARCHAR, symbol SYMBOL, side SYMBOL, price DOUBLE, @@ -2749,6 +2900,7 @@ import { connectQwpNodeClient, QwpEgressQueryError, QwpIngressAckTimeoutError, + QwpPoolResourceError, QWP_RECONNECT_EVENT_KIND, type QwpReconnectEvent, type QwpSenderError, @@ -2766,9 +2918,11 @@ function logConnection(event: QwpReconnectEvent) { const db = await connectQwpNodeClient( "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + `token=${token};` + - // query_pool_min=0: start, and ingest, even while no replica is reachable. + // append: every flush waits for a disk sync; see "Store-and-forward". "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + - "sf_durability=append;sender_pool_max=4;query_pool_min=0;query_pool_max=8;", + "sf_durability=append;sender_pool_max=4;" + + // query_pool_min=0: start, and ingest, even while no replica is reachable. + "query_pool_min=0;query_pool_max=8;", { // Queries run on replicas only, never on the primary; ingestion always // follows the primary. @@ -2817,7 +2971,7 @@ try { for (const event of events) { await sender .table("trades_sf") - .symbol("trade_id", event.tradeId) + .stringColumn("trade_id", event.tradeId) .symbol("symbol", event.symbol) .symbol("side", "buy") .doubleColumn("price", event.price) @@ -2830,31 +2984,40 @@ try { if (!(error instanceof QwpIngressAckTimeoutError)) throw error; console.warn("ACK timed out; rows remain in sf_dir for replay after close"); } finally { - await sender.close(); + // After a terminal rejection, close() rejects with the same failure. + // Log it so that it does not replace the error thrown above. + await sender.close().catch((error) => console.error("close failed:", error)); } - // Querying: rows may not be visible yet, see "Read-after-write". While no - // replica is reachable, borrowQuery() rejects with QwpPoolResourceError. - const lease = await db.borrowQuery(); + // Querying: rows may not be visible yet, see "Read-after-write". try { - const query = await lease.query( - "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + - "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", - { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, - ); - const recentPrices: (readonly unknown[])[] = []; - for await (const batch of query) { - // A failover re-executes the query from sequence 0: drop the partial result. - if (batch.batchSequence === 0n) recentPrices.length = 0; - for (const row of batch.rows()) recentPrices.push(row); + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + + "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", + { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, + ); + const recentPrices: (readonly unknown[])[] = []; + for await (const batch of query) { + // A failover restarts the result at sequence 0: drop partial rows. + if (batch.batchSequence === 0n) recentPrices.length = 0; + for (const row of batch.rows()) recentPrices.push(row); + } + await query.completion; + console.log(recentPrices); + } finally { + await lease.close(); } - await query.completion; - console.log(recentPrices); } catch (error) { - if (!(error instanceof QwpEgressQueryError)) throw error; - console.error(`query failed: status=${error.status} ${error.message}`); - } finally { - await lease.close(); + if (error instanceof QwpPoolResourceError) { + // No replica was reachable within the failover budget. + console.warn("no replica available for queries:", error.cause); + } else if (error instanceof QwpEgressQueryError) { + console.error(`query failed: status=${error.status} ${error.message}`); + } else { + throw error; + } } } finally { await db.close(); diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 6eda395eea..46c2df520f 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -31,8 +31,8 @@ Pick the path that matches your environment. The first-party libraries for **Java, Python, Go, Rust, Node.js, C & C++, and .NET** are the recommended way to talk to QuestDB. They speak the **QuestDB Wire Protocol (QWP)** and unify ingest and query under one -client configuration. The client may use separate connections for ingestion -and queries, as the Node.js library does. +client configuration. Ingestion and queries run over separate WebSocket +connections, which the clients' pools manage for you. ### QWP support @@ -54,13 +54,15 @@ Highlights: - **Binary on the wire** — roughly half the size of ILP or HTTP. - **Streaming both directions** — sustained 800 MiB/s ingress, up to 2.5 GiB/s egress on a single connection. -- **Automatic failover** — ingress and egress fail over without application - intervention. +- **Automatic failover** — ingress and egress reconnect and fail over without + application intervention. A query that fails over restarts from its first + row, so code that accumulates rows must reset them; see each client's page. - **Store-and-forward** — survives server outages, including full server destruction. Sub-200 ns offload latency. - **One configuration** — a single [connect string](/docs/connect/clients/connect-string/) drives every - option, portable across all languages. + option, with the same keys in every language. A few defaults and behaviors + differ per client, as the connect string reference notes. - **Schema-flexible** — automatic table creation and on-the-fly column additions. @@ -107,8 +109,8 @@ covering the WebSocket variants for ingress and egress. Read these if you are embedding QuestDB connectivity into an existing framework. QWP also has a UDP transport for fire-and-forget metrics, supported by the -Java, Rust, C, C++ and Node.js clients via the `udp` connect-string schema. It -is configured through the [`qwp.udp.*` server +Java, Python, Rust, C, C++ and Node.js clients via the `udp` connect-string +schema. It is configured through the [`qwp.udp.*` server settings](/docs/configuration/qwp/#udp-receiver) and is disabled by default; there is no separate byte-level specification page for it. diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index 7a02721ac4..e46db0e4e7 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -76,8 +76,9 @@ Why each line matters: It bounds only the **blocking** initial connect (`initial_connect_retry=on` / `sync`). Once a sender is running, the reconnect loop never consults it and retries a transport outage forever. Setting a large value here does nothing for -a running producer. The Node.js client is the exception: a sender with neither -`sf_dir` nor `initial_connect_retry=async` applies it to every outage. See +a running producer. The exception is a Node.js sender with neither `sf_dir` +nor background replay (`initial_connect_retry=async`, or pooled +`lazy_connect=on`), which applies it to every outage. See [Reconnect and outage handling](#reconnect-and-outage-handling). ::: @@ -202,14 +203,14 @@ callers block up to `acquire_timeout_ms` then throw. | `sf_durability` | `memory` (also supports `periodic`; Node.js also `append`) | | `sf_sync_interval_millis` | `5000` (requires `sf_durability=periodic`) | | `sf_append_deadline_millis` | `30000` | -| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** (Node.js memory mode: every outage) | +| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** (a Node.js sender without `sf_dir` or background replay: every outage) | | `reconnect_initial_backoff_millis` | `100` | | `reconnect_max_backoff_millis` | `5000` | | `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/Node.js) | -| `connect_timeout` | unset — per-endpoint TCP connect bound, must be `> 0` | +| `connect_timeout` | unset (Node.js: `15000`) — per-endpoint TCP connect bound, must be `> 0` | | `auth_timeout_ms` | `15000` | | `max_frame_rejections` | `4` | -| `poison_min_escalation_window_millis` | `5000` | +| `poison_min_escalation_window_millis` | `5000` (Node.js: `300000`) | ### Query client @@ -226,7 +227,8 @@ callers block up to `acquire_timeout_ms` then throw. There is no "retry forever" setting to look for on the reconnect keys — a running sender already does. `reconnect_max_duration_millis` applies only to a -blocking initial connect, except on a Node.js sender in memory mode; see +blocking initial connect, except on a Node.js sender with neither `sf_dir` +nor background replay; see [Reconnect and outage handling](#reconnect-and-outage-handling). --- @@ -299,7 +301,7 @@ here. | 2 | `reconnect_max_duration_millis` is named as if it governs reconnection, but a running sender never consults it — it bounds only the blocking initial connect. | Candidate (naming) | | 3 | In Java, `failover` sounds like it covers startup but only affects post-connect query `execute()`. Java queries have no async/lazy initial connect. For the Node.js exception, see the [mental model](#mental-model). | Candidate | | 4 | No first-class write-only facade: a write-only user must still supply a query config and remember `query_pool_min=0`, or use `lazy_connect=true`. | Candidate | -| 5 | A single endpoint returning `401`/`403` normally aborts the whole endpoint walk, even if other endpoints would accept the credentials. See the [Node.js authentication recovery exception](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). | Intended (documented), revisit | +| 5 | A single endpoint returning `401`/`403` aborts the whole endpoint walk, even if other endpoints would accept the credentials. Some clients retry after a sender's first connection; see [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). | Intended (documented), revisit | | 6 | Query `serverInfoTimeoutMs` has no config key, so a facade query client cannot tune it. | Candidate | | 7 | The simplest API (`fromConfig` + async) has the worst error visibility — terminal async failures surface only on later producer calls or at `close()`. | Candidate | | 8 | `SenderProgressHandler` has no builder setter on either surface; it must be installed post-construction via `QwpWebSocketSender.setProgressHandler`. | Candidate | @@ -360,34 +362,37 @@ With `initial_connect_retry=async`: - Terminal errors go to a configured `SenderErrorHandler`; without one they surface on later producer calls or at close-time. -A sender in async mode does not give up because time passed. Terminal -conditions include authentication rejection, a durable-ack capability -mismatch, and poison-frame escalation, subject to the Node.js exceptions -below. Producer calls can also fail when the buffer reaches -`sf_max_total_bytes` and exhausts `sf_append_deadline_millis` on `append()`. - -Node.js differs on authentication recovery: initial rejection is terminal, but -regular senders with `sf_dir` or background memory replay -(`initial_connect_retry=async` or `lazy_connect=on`) retry authentication -rejections after their first successful connection. See +A sender in async mode does not give up because time passed. What stops it is +a terminal condition: poison-frame escalation at any time, and an +authentication rejection or durable-ack capability mismatch before its first +successful connection. After that first connection, the Java reference client +retries authentication and durable-ack rejections instead, so a credential or +capability change on the cluster cannot stop the producer. The Rust, C, C++, +Python, Go, and .NET clients treat them as terminal. Producer calls can also +fail when the buffer reaches `sf_max_total_bytes` and exhausts +`sf_append_deadline_millis` on `append()`. + +The Node.js client retries authentication rejections after the first +connection only in senders with `sf_dir` or background memory replay +(`initial_connect_retry=async` or `lazy_connect=on`); see [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). - -For unsupported durable acknowledgement, Node.js senders with -`initial_connect_retry=async` or `lazy_connect=on` keep retrying and emit -`durable-ack-unavailable`, rather than becoming terminal. Store-and-forward -senders do the same after their first successful connection. Monitor these +For unsupported durable acknowledgement, its senders with +`initial_connect_retry=async` or `lazy_connect=on` keep retrying from startup +and emit `durable-ack-unavailable`, and store-and-forward senders do the same +after their first successful connection. Monitor these [connection events](/docs/connect/clients/nodejs/#connection-events) and buffer usage; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). ### Reconnect and outage handling -**A running sender retries a transport outage indefinitely.** There is no -wall-clock give-up and no budget-exhaustion event. Backoff grows from +**A running sender retries a transport outage indefinitely**, except the +Node.js senders described in the note below. There is no wall-clock give-up +and no budget-exhaustion event. Backoff grows from `reconnect_initial_backoff_millis` to `reconnect_max_backoff_millis` and stays there; the loop rotates through the endpoints in `addr` as it goes. -`reconnect_max_duration_millis` has exactly two consumers, neither of which is -the steady-state loop: +In the Java reference client, `reconnect_max_duration_millis` has exactly two +consumers, neither of which is the steady-state loop: 1. The **blocking sync initial connect** (`initial_connect_retry=on` / `sync`), which gives up and throws when the budget expires. @@ -404,9 +409,11 @@ has outlasted your configuration. :::note Alignment This is the behaviour of the Java reference client and the .NET client. Other -clients are aligned to it, except the Node.js client in memory mode: a -sender with neither `sf_dir` nor `initial_connect_retry=async` gives up after -`reconnect_max_duration_millis`. If you are implementing a new client, the contract +clients are aligned to it, except a Node.js sender with neither `sf_dir` nor +background replay (`initial_connect_retry=async`, or pooled +`lazy_connect=on`), which gives up after `reconnect_max_duration_millis`; see +the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). If you +are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, and apply back-pressure to the producer rather than dropping data. @@ -421,8 +428,8 @@ and apply back-pressure to the producer rather than dropping data. | TLS session/certificate failure | transport error; try next endpoint | | HTTP upgrade timeout / non-auth transport error | try next endpoint | | `421` with `X-QuestDB-Role: REPLICA` | role reject; try next endpoint | -| `401` / `403` auth failure | **terminal**; do not try later endpoints, except during [Node.js background sender recovery](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | -| durable-ack requested but unsupported | terminal mismatch, except for [Node.js background retries](/docs/connect/clients/nodejs/#durable-acknowledgement) | +| `401` / `403` auth failure | never try later endpoints; **terminal** before the first successful connection, then client-specific: [Java and some Node.js senders retry](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | +| durable-ack requested but unsupported | terminal mismatch, except that Java senders retry after their first successful connection, and Node.js senders retry in background modes, or after their first connection with `sf_dir` ([details](/docs/connect/clients/nodejs/#durable-acknowledgement)) | | successful write upgrade | bind this endpoint | | all endpoints fail transport | throw / retry per initial/reconnect mode | | all endpoints role-reject as replicas | `QwpRoleMismatchException` | diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index 55c9d78f3f..df07504d14 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1127,11 +1127,15 @@ The client uses double-buffered microbatches: | Byte size | disabled | | Time since first row | 100 ms | +The Node.js client measures the interval from its last flush, or from sender +creation, instead of from the first buffered row. + ### Failover and high availability Ingress senders use a reconnect loop regardless of whether store-and-forward -is configured. The two storage modes share identical failover semantics; they -differ only in where unacknowledged data lives: +is configured. The two storage modes share the same failover semantics, apart +from the Node.js memory-mode budget in the table below; they differ only in +where unacknowledged data lives: - **`sf_dir` set** (store-and-forward): segments are memory-mapped files under `sf_dir`. Unacknowledged data survives sender restarts and is replayed by @@ -1147,7 +1151,7 @@ section of the connect string reference: | Key | Default | Description | |----------------------------------|-----------|-------------------------------------------| -| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in memory mode. | +| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`). | | `reconnect_initial_backoff_millis` | `100` | First post-failure sleep. | | `reconnect_max_backoff_millis` | `5000` | Cap on per-attempt sleep. | | `initial_connect_retry` | `off` | Retry on first connect (`on`, `sync`, `async`). | @@ -1161,11 +1165,12 @@ Key behaviors: The Node.js client is the exception: it applies `zone=` and `target=` to ingress too, so `target=replica` in a shared connect string stops its ingestion. -- **Authentication rejection (`401`/`403`) is normally terminal.** After a - successful connection, a regular Node.js sender with `sf_dir` or background - memory replay (`initial_connect_retry=async` or `lazy_connect=on`) instead - retries it indefinitely. Initial authentication rejection remains terminal; - see the [authentication recovery exception](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). +- **Authentication rejection (`401`/`403`) never moves to another host.** It + is terminal before a sender's first successful connection. After that, the + Java client retries it indefinitely, the Node.js client does so for senders + with `sf_dir` or background memory replay (`initial_connect_retry=async` or + `lazy_connect=on`), and other clients stop; see + [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). - **`421 + X-QuestDB-Role`** is a role reject: transient if the role is `PRIMARY_CATCHUP`, topology-level otherwise. - **All other upgrade errors are transient** and feed into the reconnect loop, diff --git a/documentation/getting-started/quick-start.mdx b/documentation/getting-started/quick-start.mdx index 6a9bbee517..9608b63d36 100644 --- a/documentation/getting-started/quick-start.mdx +++ b/documentation/getting-started/quick-start.mdx @@ -269,7 +269,7 @@ Now... Time to really blast-off. 🚀 Next up: Bring your data - the _life blood_ of any database. -Choose from one of our premium ingest-only language clients: +Choose a first-party client library for ingestion and queries: diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index 85650a8998..1844d5663c 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -121,7 +121,10 @@ The `target=` key controls which server role the client is willing to bind to: up to its predecessor's WAL — the client treats it as transient and retries the same host (with a fresh round, no exponential backoff) until it becomes a full `PRIMARY`. On an ingress sender this retry has no deadline; the producer -is bounded by buffer capacity rather than by elapsed time. +is bounded by buffer capacity rather than by elapsed time. The exception is a +Node.js sender with neither `sf_dir` nor background replay +(`initial_connect_retry=async`, or pooled `lazy_connect=on`), which stops after +`reconnect_max_duration_millis`. A `421 Misdirected Request` response **without** an `X-QuestDB-Role` header is treated as a generic transport error, not a role reject — the client walks @@ -139,17 +142,20 @@ very different goals. The ingress reconnect loop sits inside the store-and-forward I/O thread. It runs continuously in the background, retrying through outages while the -producer keeps appending to the local buffer. There is no wall-clock give-up: -the loop retries an outage of any length, and what bounds your tolerance is -buffer capacity (`sf_max_total_bytes` and disk), not a timer. +producer keeps appending to the local buffer. There is no wall-clock give-up, +except in one Node.js mode described below: the loop retries an outage of any +length, and what bounds your tolerance is buffer capacity +(`sf_max_total_bytes` and disk), not a timer. - Initial backoff: `100 ms` - Maximum backoff: `5 s` - Per-outage budget: **none**. `reconnect_max_duration_millis` bounds only the blocking sync initial connect, and the running loop never consults it. - The Node.js client is the exception: in memory mode, a sender without - `initial_connect_retry=async` gives up after `reconnect_max_duration_millis` - and fails with `QwpReconnectExhaustedError`. + The exception is a Node.js sender with neither `sf_dir` nor background + replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`), which + gives up after `reconnect_max_duration_millis` and fails with + `QwpReconnectExhaustedError`; see the + [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps reconnect storms when many producers share a cluster - Inter-host pause within a round: **none** — the client walks the full @@ -191,7 +197,7 @@ will not help. | Condition | Why terminal | |---|---| -| HTTP `401` / `403` on upgrade | Credentials are assumed to be cluster-wide. See the [Node.js sender recovery exception](#authentication-is-cluster-wide). | +| HTTP `401` / `403` on upgrade | Credentials are assumed to be cluster-wide. Some clients retry after a sender's first successful connection; see [Authentication is cluster-wide](#authentication-is-cluster-wide). | | Server-status reject (SF) | Application-layer reject; replay reproduces the same response. | ### Topology — handled inside the round @@ -218,7 +224,9 @@ and walks to the next host. When a round exhausts with transient errors, the client sleeps for the backoff interval and starts the next round. On the ingress sender the rounds -continue indefinitely; on the egress query client they are bounded by +continue indefinitely, apart from the Node.js exception described under +[Ingress (writes)](#ingress-writes); on the egress query client they are +bounded by `failover_max_attempts` and `failover_max_duration_ms`, which apply per `execute()`. @@ -237,18 +245,24 @@ when at least one peer is healthy." ## Authentication is cluster-wide -A `401` or `403` on the HTTP upgrade is normally terminal: the client does not -retry other hosts. Credentials are assumed to be configured identically across -the cluster, so trying another node would repeat the rejection. - -The Node.js client makes an exception for a regular ingress sender that has -already connected successfully and uses `sf_dir` or background memory replay -(`initial_connect_retry=async` or `lazy_connect=on`). It retries authentication -rejections indefinitely so buffered data can drain after server-side -authentication is restored. Initial authentication rejection remains terminal, -including in these modes. The exception does not apply to query connections, -ordinary memory-only senders, or orphan drainers. See -[Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors). +A `401` or `403` on the HTTP upgrade does not send the client to another host: +credentials are assumed to be configured identically across the cluster, so +another node would repeat the rejection. Whether the rejection is terminal +depends on the client and on when it arrives: + +| Client | Before a sender's first successful connection | After it | +|---|---|---| +| Java | Terminal | Retried indefinitely | +| Node.js | Terminal | Retried indefinitely by senders with `sf_dir` or background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`). Terminal for other senders | +| Rust, C, C++, Python, Go, .NET | Terminal | Terminal | + +Query connections and orphan drainers treat the rejection as terminal in every +client. A sender that retries keeps buffering until authentication succeeds +again, bounded by its buffer capacity, so a credential change on the cluster +does not stop the producer. Watch its connection events rather than waiting +for an error. See +[Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors) +for the Node.js rules. Per-host credentials are outside the failover model. Use a separate connect string for each credential. diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index ffa360d4c3..8a76aafaf4 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -15,9 +15,10 @@ first. ## Common keys `addr` and `auth_timeout_ms` apply to every WS / WSS / HTTP / HTTPS client. -`zone` is accepted everywhere but only takes effect on egress; `target` is an -egress-only key and is rejected as an unknown key on an ingress connect string. -The Node.js client is the exception: it applies both keys to ingress too. +`zone` and `target` are accepted everywhere but only take effect on egress: +ingress parsers accept and ignore them, so one connect string can serve both +directions. The Node.js client is the exception: it applies both keys to +ingress too. They are documented in full on the [connect-string reference](/docs/connect/clients/connect-string#failover-keys); the table below summarises the failover-relevant subset. @@ -26,8 +27,8 @@ the table below summarises the failover-relevant subset. |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | | `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Rejected as an unknown key on an ingress connect string. The [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | -| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. The Node.js client bounds DNS and TCP/TLS separately with `connect_timeout` (15 s by default); other clients may use the OS default. Lower it for faster upgrade failure detection; tune the connect timeout separately. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Accepted and ignored on an ingress connect string. The [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. The Node.js client bounds DNS and TCP/TLS separately with `connect_timeout` (15 s by default); other clients may use the OS default. On Node.js, `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. Lower it for faster upgrade failure detection; tune the connect timeout separately. | `addr` syntax — both of these are equivalent and produce the same three-peer list: @@ -48,7 +49,7 @@ for the full list. The failover-relevant keys are: | Key | Type | Default | Notes | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender with neither `sf_dir` nor `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | @@ -62,7 +63,7 @@ network), and retrying for five minutes only hides it. | Value | Behaviour | |---|---| | `off` (default; alias `false`) | First-connect failure is terminal. The producer's call to build the sender throws immediately. | -| `on` (aliases `sync`, `true`) | First-connect failures are retried on the caller's thread. The constructor blocks until it connects or `reconnect_max_duration_millis` expires — this is the **only** place that key applies. Once the sender is running, reconnection is unbounded. A Node.js sender in memory mode is the exception; see the `reconnect_max_duration_millis` row above. | +| `on` (aliases `sync`, `true`) | First-connect failures are retried on the caller's thread. The constructor blocks until it connects or `reconnect_max_duration_millis` expires — this is the **only** place that key applies. Once the sender is running, reconnection is unbounded, except for the Node.js senders described in the `reconnect_max_duration_millis` row above. | | `async` | The constructor returns immediately; the background I/O thread drives the reconnect loop. The producer experiences backpressure if it tries to publish before the connection comes up. Intended for unattended producers where the SF directory may already carry segments from a prior process and the server may come up later. | ## Egress (query) diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index b636029303..6bb17c787a 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -19,7 +19,9 @@ arrive asynchronously. A network outage or a server restart leaves your producer code unaffected — the I/O thread quietly reconnects and replays what remains. In SF mode, even a crash of the sender process itself loses no unacked data: the next sender on the slot recovers it from disk and -replays it. +replays it. The one exception is a Node.js sender with neither `sf_dir` nor +background replay, whose `flush()` waits for the reconnect during an outage; +see [Reconnect and replay](#reconnect-and-replay). ## Two modes @@ -38,10 +40,10 @@ SF runs in either of two modes selected by the connect string: Both modes share the same reconnect loop, the same backoff and retry budgets, and the same on-the-wire behaviour. The only difference is -where unacked data lives. The Node.js client is the exception: ordinary -memory mode gives up after `reconnect_max_duration_millis` (5 minutes by -default). Background memory mode (`initial_connect_retry=async`, also selected -by pooled `lazy_connect=on`) and disk-backed SF retry indefinitely. See the +where unacked data lives. The Node.js client is the exception: a sender with +neither `sf_dir` nor background replay (`initial_connect_retry=async`, or +pooled `lazy_connect=on`) gives up after `reconnect_max_duration_millis` +(5 minutes by default), while its other senders retry indefinitely. See the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). ## What "frame" means here @@ -134,10 +136,14 @@ object store** (S3, Azure Blob, GCS, or NFS). advances to the highest covered wireSeq. - The client requires an `X-QWP-Durable-Ack: enabled` echo on the upgrade response and rejects a connection without it, rather than waiting for ack - frames it cannot receive. This is normally terminal. Node.js background - senders instead retry and emit `durable-ack-unavailable`; see - [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement) - for the retry modes and monitoring requirements. + frames it cannot receive. In most clients the rejection is terminal. The + Java client retries it after a sender's first successful connection, so a + capability change on the cluster cannot stop the producer. Node.js background senders + (`initial_connect_retry=async`, or pooled `lazy_connect=on`) retry it from + startup, and Node.js store-and-forward senders after their first + connection, emitting `durable-ack-unavailable` events; see + [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). + A retrying sender keeps buffering, so monitor it. Durable-ack mode is the right choice when "data is in the object store" is the durability bar, but it has two costs: a longer time-to-trim (so @@ -156,8 +162,9 @@ the reconnect loop documented in The producer is **not notified**: it keeps publishing into the substrate, subject to available capacity (see [Backpressure](#backpressure)). -On Node.js, only ordinary memory mode waits for reconnect in `flush()`, up to -`reconnect_max_duration_millis`. Background memory mode, enabled by +On Node.js, only a sender with neither `sf_dir` nor background replay waits +for the reconnect in `flush()`, up to `reconnect_max_duration_millis`. +Background memory mode, enabled by `initial_connect_retry=async` or pooled `lazy_connect=on`, keeps accepting batches into the memory replay queue until capacity is exhausted and retries indefinitely. See the [three Node.js flushing modes](/docs/connect/clients/nodejs/#flushing). @@ -218,8 +225,10 @@ fires, a `WARN` is logged and: - in **memory mode**, the un-acked tail is lost. On the Node.js client, a standalone sender's `close()` rejects with -`QwpSenderCloseTimeoutError` instead of logging, and the pooled client's -`db.close()` reports the timeout to its `onError` callback. +`QwpSenderCloseTimeoutError` instead of logging. The pooled client's +`db.close()` resolves, and usually reports the timeout to +`ingressSession.onError` as a non-terminal `QwpIngressAckTimeoutError`; that +report is best-effort. Setting `close_flush_timeout_millis=0` (or `-1`) skips the drain wait entirely — useful for fast shutdown paths where you do not want to block. @@ -331,24 +340,35 @@ shared `sf_dir`, blindly draining unknown slots may be surprising. Not every server response is an OK. A rejected batch is **not** silently dropped and trimmed: the client either retries it or reports a terminal error. -The defaults below apply to the Node.js QWP client; consult the +There is no drop policy. The table lists the built-in default policy of each +category in the Java reference client and the Node.js client; see the [connect-string error policies](/docs/connect/clients/connect-string/#error-handling) -for the shared vocabulary and per-client override support. +for the shared vocabulary and which clients let you override the defaults. -| Category | Node.js default | Meaning | +| Category | Default policy | Meaning | |---|---|---| | `SCHEMA_MISMATCH` | `terminal` | The schema does not match. The sender stops; in SF mode, the rejected bytes remain in the journal for inspection. | -| `WRITE_ERROR` | `retriable` | A write failed (for example, a temporary storage problem); reconnect and replay. | | `PARSE_ERROR` | `terminal` | Malformed payload; replaying identical bytes cannot help. | -| `INTERNAL_ERROR` | `retriable` | Retry after an unexpected server-side failure. | -| `SECURITY_ERROR` | `terminal` | Authentication or authorization failed. | +| `SECURITY_ERROR` | `terminal` | Authorization failed, for example an ACL denial on a writable node. | | `PROTOCOL_VIOLATION` | `terminal` (forced) | Protocol failure; stop and report it. | - -The Node.js client delivers asynchronous rejections to `onSenderError`; its -default handler logs them. Other clients may use a bounded error inbox that -drops the oldest notification on overflow. Check the client-specific error -handling before relying on a policy override: the Node.js `on_*_error` keys -are currently accepted but do not change these defaults. +| `WRITE_ERROR` | `retriable` | A write failed, for example under temporary storage pressure; reconnect and replay. | +| `INTERNAL_ERROR` | `retriable` | Retry after an unexpected server-side failure. | +| `DICTIONARY_GAP` | `retriable` | The connection is missing symbol dictionary entries; resend them and replay. | +| `NOT_WRITABLE` | `retriable_other` | The node cannot accept writes, for example a replica; replay on another endpoint. | +| `UNKNOWN` | `retriable` (forced) | A status the client does not know, for example from a newer server; retry rather than stop. | + +A batch that keeps being rejected without progress escalates to a terminal +error through the poison-frame detector (`max_frame_rejections`). The Java and +Node.js clients also report a client-side `DATA_LOSS` category, with the +`abandoned` policy, when they set aside a corrupt store-and-forward journal. + +Rejections are delivered asynchronously through a bounded error inbox +(`error_inbox_capacity`, default `256`) that drops the oldest notification on +overflow, to the application's error handler, such as `onSenderError` on +Node.js. The default handler logs every rejection, because a silent handler +would hide data loss. The Node.js client reports categories as lowercase, +hyphenated `error.category` strings, such as `schema-mismatch`, and accepts +the `on_*_error` keys without applying them. ## Next steps diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 1e238cbdf4..f9cf444ded 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -48,7 +48,7 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender with neither `sf_dir` nor `initial_connect_retry=async` applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | @@ -64,7 +64,7 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| -| `request_durable_ack` | bool | `off` | Opt-in via the upgrade header `X-QWP-Request-Durable-Ack: true`. Trim is then driven by `STATUS_DURABLE_ACK` frames only; OK frames no longer advance the trim watermark. A missing `X-QWP-Durable-Ack: enabled` echo is terminal except for [Node.js background retries](/docs/connect/clients/nodejs/#durable-acknowledgement). WebSocket transports only. | +| `request_durable_ack` | bool | `off` | Opt-in via the upgrade header `X-QWP-Request-Durable-Ack: true`. Trim is then driven by `STATUS_DURABLE_ACK` frames only; OK frames no longer advance the trim watermark. A missing `X-QWP-Durable-Ack: enabled` echo is terminal in most clients; Java and some Node.js senders keep retrying (see [Concepts](/docs/high-availability/store-and-forward/concepts/#trim-how-unacked-data-is-reclaimed)). WebSocket transports only. | | `durable_ack_keepalive_interval_millis` | int (ms) | `200` | Cadence of WebSocket PING the I/O loop sends while there are pending durable confirmations and the producer is idle. `0` or negative disables. | ## Error-handling keys @@ -72,12 +72,13 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| | `error_inbox_capacity` | int (≥16) | `256` | Bounded SPSC queue capacity for async error notifications. Overflow drops the oldest entry and increments `getDroppedErrorNotifications`. | -| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | All clients accept these keys, but Node.js and Java currently ignore them; .NET applies them. There is no `DROP_AND_CONTINUE` policy. See [Error handling](/docs/connect/clients/connect-string/#error-handling). | +| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | All clients accept these keys. Go and .NET apply them; Java, Node.js, Rust, C, C++, and Python currently ignore them. There is no `DROP_AND_CONTINUE` policy. See [Error handling](/docs/connect/clients/connect-string/#error-handling). | -The Node.js defaults are documented in +The per-category defaults are documented in [Concepts § Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). -`PROTOCOL_VIOLATION` is always terminal; Node.js treats an unknown server -status as retriable rather than silently dropping the batch. +`PROTOCOL_VIOLATION` is always terminal and `UNKNOWN` always retriable, so a +status from a newer server leads to a retry rather than a silently dropped +batch. ## Other relevant keys diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index 1810f292b3..83fcafd497 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -56,26 +56,18 @@ incompatible. :::caution Node.js client -The Node.js client does not use an OS lock. It locks a slot by creating a -`.lock.owner` directory inside it, which records the owner's host name and -process ID, and keeps `.lock` and `.lock.pid` only for compatibility. A -crashed Node.js sender therefore leaves the slot locked: a new sender takes it -over automatically only on the same host, once the recorded process ID is no -longer in use. That often fails in containers, where the application usually -runs as process ID 1 and a replacement container has a new host name. -Otherwise, verify that the previous owner has exited and no process is using -that slot before removing its stale `//.lock.owner` directory. -Here `` is ``, or `-` for a pooled Node.js sender. - -If startup still reports `QwpReplayStoreLockedError`, inspect -`/.slot-locks/.lock.owner` too. This short-lived guard can survive -a crash during lock acquisition or quarantine. Remove only that specific -owner directory after verifying its owner has exited, never the shared -`.slot-locks` directory or another slot's locks. +The Node.js client does not use an OS lock. It locks a slot with a +`.lock.owner` directory that records the owner's host name and process ID, and +keeps `.lock` and `.lock.pid` only for compatibility. After a crash, a new +Node.js sender takes the slot over automatically only on the same host, once +the recorded process ID is no longer in use. In containers that usually fails, +because the application runs as process ID 1 and a replacement container has a +new host name, and the new sender reports `QwpReplayStoreLockedError`. The +[Node.js client](/docs/connect/clients/nodejs/#store-and-forward) describes how +to remove a stale lock safely. Node.js and other clients do not see each other's locks, so never let them use -the same `sf_dir` at the same time. See the -[Node.js client](/docs/connect/clients/nodejs/#store-and-forward). +the same `sf_dir` at the same time. ::: @@ -117,8 +109,9 @@ when the new one comes up. Solutions: - Stop the previous process. For clients using OS locks, the kernel releases the lock on exit (even after `kill -9`). A killed Node.js sender can leave - stale owner directories behind: verify the old owner is gone before removing - the specific directories described in [`.lock` and `.lock.pid`](#lock-and-lockpid). + a stale `.lock.owner` directory behind: verify that the old owner is gone + before removing it, as the + [Node.js client](/docs/connect/clients/nodejs/#store-and-forward) describes. - Use a deployment unit that orders shutdown before startup. - For containerised deployments, set `sender_id` from a per-pod stable identity so two pods with the same template name don't collide. @@ -198,7 +191,8 @@ records every producer thread that hit the cap. When an SF-mode sender opens, it runs this sequence: -1. Acquire `//.lock`. Fail loudly on contention. +1. Acquire the slot lock, `//.lock` (the Node.js client + uses a `.lock.owner` directory instead). Fail loudly on contention. 2. Scan every `*.sfa` file: - Validate magic, version, header. - Walk frames forward verifying each CRC32C-Castagnoli. @@ -222,9 +216,9 @@ fresh start: no segments, no replay. | Symptom | Likely cause | Operator action | |---|---|---| -| "Slot held by PID ``" or `QwpReplayStoreLockedError` (Node.js) | Another process holds the slot, or a Node.js `.lock.owner` is stale after a crash. | Stop the duplicate. OS locks release on exit; for Node.js verify the owner is gone before removing `.lock.owner` (see [`.lock` and `.lock.pid`](#lock-and-lockpid)). | +| "Slot held by PID ``" or `QwpReplayStoreLockedError` (Node.js) | Another process holds the slot, or a Node.js `.lock.owner` is stale after a crash. | Stop the duplicate. OS locks release on exit; for Node.js verify the owner is gone before removing `.lock.owner` (see the [Node.js client](/docs/connect/clients/nodejs/#store-and-forward)). | | "Gap between segments" | Corruption — a segment was deleted out of band. | Restore from backup or accept data loss; the substrate refuses to start. | -| "Watermark exceeds publishedFsn" | `.ack-watermark` is corrupt; the engine falls back to the no-watermark seed. | Logged as `WARN`. Replay will re-send the lowest segment's frames; rely on server deduplication. | +| "Watermark exceeds publishedFsn" | `.ack-watermark` is corrupt; the engine falls back to the no-watermark seed. | Logged as `WARN`. Replay will re-send the lowest segment's frames, which inserts duplicate rows unless the table has `DEDUP UPSERT KEYS`. | | Torn tail count > 0 | The previous process crashed mid-frame-write. | Informational; the CRC + zero-fill design discards the partial frame. | ## Close and shutdown @@ -233,7 +227,7 @@ fresh start: no segments, no replay. | Value | Behaviour | |---|---| -| `5000` (default) | Block up to 5 s waiting for `ackedFsn ≥ publishedFsn`. Log `WARN` on timeout; un-acked tail stays on disk (SF) or is lost (memory). | +| client default: `60000` on Java and .NET, `5000` on Rust, C, C++, Python, and Node.js | Block up to that long waiting for `ackedFsn ≥ publishedFsn`. Log `WARN` on timeout; un-acked tail stays on disk (SF) or is lost (memory). | | `0` or `-1` | Skip the drain wait. Pending data persists on disk (SF) for the next sender, or is lost (memory). | | any other positive value | That timeout in milliseconds. | diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 283a21547d..1d1f35cc44 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -57,9 +57,9 @@ Unacked frames are written to mmap'd files under Both modes share the same wire behaviour, the same failover loop, and the same connect-string keys for everything other than storage. You can switch between them without changing application code — only the connect -string. On the Node.js client, memory mode also gives up after -`reconnect_max_duration_millis` of outage, unless `initial_connect_retry=async` -is set. +string. On the Node.js client, a sender with neither `sf_dir` nor background +replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) also gives +up after `reconnect_max_duration_millis` of outage. ## Comparison at a glance @@ -102,23 +102,25 @@ GCS, or NFS). - WAL-local durability on the primary is sufficient. - You want minimum steady-state disk usage. - You are running OSS or a build that does not support durable-ack. - Opting in rejects connection attempts; see [Caveats](#caveats) for the - Node.js background-retry exception. + Opting in makes those connection attempts fail; see [Caveats](#caveats) for + the clients that keep retrying. ### Caveats - **Server support is required.** The client sends `X-QWP-Request-Durable-Ack: true` on the upgrade. The server must echo back `X-QWP-Durable-Ack: enabled`. Without it, for example on an OSS build - or an uninitialised primary, the connection attempt is rejected. This is - normally terminal, subject to the Node.js exception below. -- **Node.js background retries.** Senders with `initial_connect_retry=async` - or `lazy_connect=on` keep retrying unsupported durable acknowledgement - instead of failing initialization. A store-and-forward sender also retries - after its first successful connection. They emit `durable-ack-unavailable` + or an uninitialised primary, the connection attempt is rejected. In most + clients this is terminal. The exceptions follow. +- **Senders that keep retrying.** The Java client retries after a sender's + first successful connection. Node.js senders with + `initial_connect_retry=async` or `lazy_connect=on` retry from startup + instead of failing initialization, and Node.js store-and-forward senders + retry after their first successful connection; they emit + `durable-ack-unavailable` [connection events](/docs/connect/clients/nodejs/#connection-events). - Monitor these events and buffer usage: continued buffering can fill the - journal or memory queue even though startup succeeded. See + Monitor retrying senders and their buffer usage: continued buffering can + fill the journal or memory queue even though startup succeeded. See [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). - **Idle keepalive.** The OSS server only flushes pending durable-ack frames during inbound recv events. The client sends a WebSocket PING @@ -179,7 +181,7 @@ If you are currently using HTTP or TCP ILP ingest, the comparison is: | Server outage tolerance | Best-effort retry | None | Reconnect loop with multi-minute budget | | Multi-host failover | Yes (HTTP only) | No | Yes | | Cross-region durability ack | No | No | Yes (`request_durable_ack=on`) | -| Cluster-wide ordering | Best-effort | Best-effort | FSN-driven, server-deduplicated | +| Cluster-wide ordering | Best-effort | Best-effort | FSN-ordered; replay is at least once, so use `DEDUP UPSERT KEYS` | The transition is application-transparent — `Sender.fromConfig` accepts a `ws::` or `wss::` connect string and the public builder API is the diff --git a/documentation/query/overview.md b/documentation/query/overview.md index 469e3db7c2..c422f27212 100644 --- a/documentation/query/overview.md +++ b/documentation/query/overview.md @@ -98,9 +98,8 @@ against the demo instance. The official client libraries speak the QuestDB Wire Protocol (QWP), a binary protocol for both ingestion and querying, configured with one connection -string. A client may use separate connections for those operations: the -[Node.js client](/docs/connect/clients/nodejs/#the-connection-pool), for example, -maintains separate sender and query pools. +string. Ingestion and queries run over separate WebSocket connections, which +the clients' pools manage for you. Results stream rather than arriving in one block. The server sends batches as it produces them, so an application starts processing the head of a result @@ -109,12 +108,18 @@ has to be materialized at all. Connections can recover from transport failures. With failover enabled, a client may reconnect and re-execute an in-flight query, including on the same -host. Result rows then restart from the beginning. Some clients' materializers -discard their partial result automatically; if you process batches yourself, -reset any accumulated state on replay or handle the client's terminal error. -See [Node.js query failover](/docs/connect/clients/nodejs/#query-failover) for -an example. Re-execution can also repeat SQL writes; see -[DDL and DML statements](/docs/connect/clients/nodejs/#ddl-and-dml-statements). +host, and the result then restarts from its first row. Helpers that collect a +whole result, such as Python's `to_pandas()`, discard the partial result for +you. If you process batches yourself, reset any accumulated state when the +result restarts, or handle the client's error. Re-execution can also repeat +SQL writes such as `INSERT`. Each client describes its failover behavior: +[Java](/docs/connect/clients/java/#query-failover), +[Python](/docs/connect/clients/python/#reader-failover), +[Go](/docs/connect/clients/go/#query-failover), +[Rust](/docs/connect/clients/rust/#failover-and-errors), +[C and C++](/docs/connect/clients/c-and-cpp/#failover-retry-and-pool-lifecycle), +[.NET](/docs/connect/clients/dotnet/#failover-and-high-availability), and +[Node.js](/docs/connect/clients/nodejs/#query-failover). The Rust, C++, and Python clients hand back results as Arrow record batches. That is the native memory layout of @@ -122,8 +127,15 @@ That is the native memory layout of [Polars](/docs/integrations/data-processing/polars/), and DuckDB, so a query becomes a DataFrame with no row-by-row conversion in between. -See the [Connect overview](/docs/connect/overview/) for the per-language -guides and which clients ship QWP today. +To get started, see the querying section of your client: +[Java](/docs/connect/clients/java/#querying-with-query-and-completion), +[Python](/docs/connect/clients/python/#querying), +[Go](/docs/connect/clients/go/#querying-and-sql-execution), +[Rust](/docs/connect/clients/rust/#querying), +[C and C++](/docs/connect/clients/c-and-cpp/#querying-data), +[.NET](/docs/connect/clients/dotnet/#querying-and-sql-execution), or +[Node.js](/docs/connect/clients/nodejs/#querying). The +[Connect overview](/docs/connect/overview/) compares the clients. ## PostgreSQL From 23c26ef2b6b5c7d9f3b189e722ac820b541c6c14 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 11:50:29 +0100 Subject: [PATCH 13/25] docs(nodejs): fix data-loss and failover guidance found in review - State that a terminally rejected batch blocks store-and-forward ingestion for every table, across restarts, and that longArrayColumn() triggers it on current servers - Say which at() rejections lose staged rows, including the next borrower's rows on a failed pooled sender - Document that a locked journal stops connectQwpNodeClient() even with lazy_connect=on, and when lock cleanup may be automated - Document that batches staged before the first connection can exceed the server limit, and recommend sf_max_segment_bytes=1m - Warn that queries without a credit window buffer results in memory - Warn that QWP auto-creates tables without DEDUP; drop sf_dir from the lazy-start example - Correct the typed reconnect object's effect on the first connection - Explain that 8 fast-failing attempts end query failover in seconds, and use maxAttempts: 0 in the full example - Add a pattern for committing source offsets after acknowledgement - Move Read-after-write and the typed reconnect policy next to their prerequisites, add an error state table, and link the 4.x migration notes from the top - Qualify the connect-string recipes that share target=replica, which the Node.js client also applies to ingress --- .../connect/clients/connect-string.md | 13 +- documentation/connect/clients/nodejs.md | 592 ++++++++++++------ 2 files changed, 409 insertions(+), 196 deletions(-) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index ff00905635..1f0ab5baff 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -16,8 +16,10 @@ deviates, the affected key section and client page call it out. One `ws::` / `wss::` connect string serves both the ingress sender and the egress query client. Each direction reads the keys relevant to it and ignores keys meant only for the other direction, so the same string -configures both without edits. The *Applies to:* tag on each section below -marks which direction a key affects. +configures both without edits. The Node.js client is the exception for +`target` and `zone`, which it also applies to ingress; see +[Role filter and zone preference](#role-filter-and-zone-preference). The +*Applies to:* tag on each section below marks which direction a key affects. For legacy InfluxDB Line Protocol (ILP) transports (`http`, `https`, `tcp`, `tcps`), see the [ILP overview](/docs/connect/compatibility/ilp/overview/). @@ -144,6 +146,11 @@ wss::addr=node-a:9000,node-b:9000;sf_dir=/var/lib/myapp/qdb-sf;sender_id=ingest- wss::addr=node-a:443,node-b:443;target=replica;zone=eu-west-1a; ``` +Senders in other clients ignore `target`, so they can share this string. On the +Node.js pooled client, `target=replica` in the connect string stops ingestion; +set the role with the typed `egress` option instead, as described under +[Role filter and zone preference](#role-filter-and-zone-preference). + ### Tolerate a slow or restarting server at startup ``` @@ -169,7 +176,7 @@ caveats), follow the section links from the [Key index](#key-index). | Bearer-token credentials | both | `token` | `auth_timeout_ms` | | Multi-host failover | both | `addr=h1,h2,…` | `target`, `zone`, `reconnect_*` (ingress), `failover_*` (egress) | | Query only the primary (freshest data) | egress | `target=primary` | — | -| Query only replicas (offload primary) | egress | `target=replica` | — | +| Query only replicas (offload primary) | egress | `target=replica` | Node.js: use the typed `egress.target` option, because the key also filters ingress | | Zone-aware routing with DR last-resort | egress | `zone=` | `target` | | Tune ingest batching | ingress | — | Clients with auto-flush: `auto_flush_rows`, `auto_flush_interval`, `auto_flush_bytes` | | Disable auto-flush (manual `flush()` only) | ingress | `auto_flush=off` | — | diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 53963690ef..3b3b26ead1 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -41,10 +41,12 @@ Key capabilities: QWP clients in a few places; see [Differences from other clients](#differences-from-other-clients). -:::tip Legacy transports +:::tip Upgrading from 4.x or using ILP -The Node.js `Sender` class still speaks ILP over HTTP and TCP. This page -documents the recommended QWP path. For ILP, see +Version 5.0.0 adds QWP and changes how the existing `Sender` handles `null` +and `undefined` values; see [Upgrading from 4.x](#upgrading-from-4x). To move +existing ILP code to QWP, see [From ILP to QWP](#from-ilp-to-qwp). The +`Sender` class still speaks ILP over HTTP and TCP; for those transports, see [ILP transports (legacy)](#ilp-transports-legacy) near the end of this page. ::: @@ -170,79 +172,7 @@ timestamp column named `timestamp`. Timestamps come back as `bigint` microseconds since the Unix epoch; see [Reading result values](#reading-result-values) for every type. -### Read-after-write - -When `flush()` resolves, the client has published the rows, but QuestDB may not -have received them yet. QuestDB acknowledges a batch once it has committed it -to its write-ahead log, and applies committed rows to the table -asynchronously. A query that runs right after ingestion can therefore fail -with `table does not exist` on a first run, or succeed and return no rows. - -When your code must read its own writes, create the table first, write an event -with a unique ID, and poll for **that ID**. Pre-creating the table avoids -mistaking an unrelated SQL error for the first-write table-creation delay. Give -each query the time remaining until the deadline so a stalled query cannot -leave the poll running indefinitely: - -```typescript -import { randomUUID } from "node:crypto"; -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const ddl = await lease.query( - "CREATE TABLE IF NOT EXISTS trades_readback (" + - "timestamp TIMESTAMP, trade_id VARCHAR, symbol SYMBOL" + - ") TIMESTAMP(timestamp) PARTITION BY DAY", - ); - await ddl.completion; - - const tradeId = randomUUID(); - const sender = await db.borrowSender(); - try { - await sender - .table("trades_readback") - .stringColumn("trade_id", tradeId) - .symbol("symbol", "ETH-USD") - .at(Date.now(), "ms"); - await sender.flush(); - } finally { - await sender.close(); - } - - const deadline = Date.now() + 10_000; - let visible = false; - while (!visible) { - const remainingMs = deadline - Date.now(); - if (remainingMs <= 0) throw new Error("trade not visible in time"); - const query = await lease.query( - "SELECT trade_id FROM trades_readback WHERE trade_id = $1 LIMIT 1", - { - binds: (binds) => binds.setVarchar(0, tradeId), - timeoutMs: remainingMs, - }, - ); - for await (const batch of query) visible ||= batch.rowCount > 0; - await query.completion; - if (!visible) { - await new Promise((resolve) => - setTimeout(resolve, Math.min(100, Math.max(0, deadline - Date.now()))), - ); - } - } - console.log(`visible trade: ${tradeId}`); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -SQL errors now surface instead of being retried. Do not replace the poll with a -fixed sleep: the apply latency varies with load. +To read your own writes reliably, see [Read-after-write](#read-after-write). ## Connecting @@ -316,13 +246,12 @@ try { `Sender` is ingestion-only. It accepts the complete QWP connect-string vocabulary, and logs a warning for keys that only the pooled client can apply, -such as `query_pool_max` or `compression`. Its fluent API covers the column -methods that also exist for ILP: `symbol`, `stringColumn`, `booleanColumn`, -`floatColumn` (DOUBLE), `intColumn` (LONG), `timestampColumn`, `arrayColumn`, -`decimalColumn`, and `decimalColumnText`. For the other QuestDB types (UUID, -IPv4, DATE, INT, and more), use a [compiled writer](#compiled-object-row-writers) -through `sender.writer()`, a pooled sender, or `connectQwpNodeSender()`, which -all expose every [column method](#column-methods). +such as `query_pool_max` or `compression`. Its fluent API has only the nine +column methods that also exist for ILP, listed under +[Column methods](#column-methods). For every other QuestDB type, use its +[compiled writer](#compiled-object-row-writers) through `sender.writer()`, +which supports every type, or use a pooled sender or `connectQwpNodeSender()`, +which expose every column method. `connectQwpNodeSender()` builds a standalone `QwpSender` from typed options. Its first argument takes the full ingestion URL, and credentials as an @@ -445,32 +374,9 @@ directions, such as `agent` or `connectTimeoutMs`) and `storeAndForward` where `qwp` has the sections `webSocket`, `session` (the equivalent of `ingressSession`), `sender`, and `udp`. -:::caution A typed `reconnect` object replaces the connect-string keys - -`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces -the whole reconnect policy parsed from the `reconnect_*`, -`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and -`egressSession.reconnect` replaces the policy parsed from `failover*` keys. -Fields you leave out of the object take the built-in defaults, not the values -from the connect string. When you supply the object, for example to register -`onEvent`, set every bound you rely on in it. - -::: - -The object's fields and the connect-string keys they replace: - -| Field | Ingestion key, default | Query key, default | -|---|---|---| -| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8` | -| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | -| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | -| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | -| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | -| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | -| `onEvent` | None | None | - -`egressSession: { reconnect: false }` is the typed equivalent of -`failover=off`. +The `reconnect` objects in `ingressSession` and `egressSession` replace the +whole reconnect policy parsed from the connect string; see +[Typed reconnect policy](#typed-reconnect-policy) before you set one. @@ -689,18 +595,14 @@ When creating a new pooled connection fails, the borrow rejects with `connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set `lazy_connect=on` to start regardless: senders connect in the background and -buffer rows until QuestDB is reachable, in memory or, with `sf_dir`, in the -[store-and-forward](#store-and-forward) journal. The query pool stays empty +buffer rows in memory until QuestDB is reachable. The query pool stays empty until the first query. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; // Resolves immediately, even if QuestDB is not running yet. -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;lazy_connect=on;" + - "sf_dir=/var/lib/my-service/qdb-sf;sender_id=startup-a;sf_durability=append;", -); +const db = await connectQwpNodeClient("ws::addr=localhost:9000;lazy_connect=on;"); try { const sender = await db.borrowSender(); try { @@ -719,18 +621,31 @@ try { } ``` -Use a writable, persistent `sf_dir` and reuse the same `sender_id` after a -restart so the example's rows survive shutdown while QuestDB is down. Without -`sf_dir`, the example would discard them when `db.close()` finishes. +Rows buffered while QuestDB is down exist only in memory. They are lost if the +client closes before QuestDB becomes reachable, as `db.close()` does at the end +of this example; see [Closing the pooled client](#closing-the-pooled-client). +To keep them across a shutdown or restart, add a +[store-and-forward](#store-and-forward) journal with `sf_dir`. Replay from the +journal is at least once, so write to a deduplicated table as described there. `lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, and rejects an explicit conflicting value. Setting `initial_connect_retry=async` without `lazy_connect` is not enough: the query pool still connects at startup, so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. -Without `sf_dir`, the buffered rows exist only in memory, and they are lost if -the client closes before QuestDB becomes reachable; see -[Closing the pooled client](#closing-the-pooled-client). + +Two caveats apply to a lazy start: + +- **A locked journal still stops it.** With `sf_dir`, the client opens its + journal at startup. If another process holds the journal, or a crashed + process left a stale lock, `connectQwpNodeClient()` rejects with + `QwpPoolResourceError` whose `cause` is `QwpReplayStoreLockedError`, with or + without `lazy_connect=on`. See Lock recovery under + [Store-and-forward](#store-and-forward). +- **Batches are not checked against the server's size limit.** Until a sender + has connected once, it cannot know the limit. If a batch can exceed about + 1 MiB, also set `sf_max_segment_bytes`; see Oversized rows under + [Flushing](#flushing). ### Closing the pooled client @@ -822,10 +737,23 @@ method called before that throws `table name must be set before adding columns`. drops every row staged since the last flush. An awaited `at()` or `atNow()` can also reject because an auto-flush failed -after the row was completed. This does not mean the row was discarded: -completed rows can remain staged for a later `flush()` or `close()`, or be -queued for replay. Do not blindly resubmit a row because its `at()` promise -rejected. See [Flushing](#flushing) and [Ingestion errors](#ingestion-errors). +after the row was completed. What happens to the completed rows depends on the +error: + +- `QwpReplayRejectedError`, `QwpIngressNackError`, `QwpReconnectExhaustedError`, + or `QwpReplayStoreError`: the sender has failed permanently, and every row + still staged on it is lost when it closes. Close the sender, borrow or create + a new one, and write those rows again. With `sf_dir`, fix the cause of a + terminal rejection first, because the rejected batch blocks the journal (see + [Ingestion errors](#ingestion-errors)). A full SYMBOL dictionary also fails + the sender permanently; see [Column methods](#column-methods). +- `QwpBatchTooLargeError`, `QwpMemoryReplayAppendTimeoutError`, or + `QwpReplayStoreAppendTimeoutError`: the completed rows stay staged for a + later `flush()` or `close()`, and the sender stays usable. Do not resubmit + them. + +These classes have no common base class, so test for them by name. See +[Flushing](#flushing) and [Ingestion errors](#ingestion-errors). ### Column methods @@ -858,7 +786,7 @@ does not exist yet: | `decimal128Column(name, unscaled, scale)` | DECIMAL(38, scale) | Unscaled `bigint`, scale up to 38 | | `decimal256Column(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, scale up to 76 | | `arrayColumn(name, value)` | DOUBLE[], DOUBLE[][], ... | Nested `number` arrays of uniform shape, 1 to 32 dimensions | -| `longArrayColumn(name, value)` | LONG[] | Encoded for protocol parity, but current QuestDB servers reject LONG array ingestion | +| `longArrayColumn(name, value)` | LONG[] | Encoded for protocol parity. Current QuestDB servers reject LONG arrays terminally; see [Arrays](#arrays) | Names that differ from what you might expect: @@ -885,7 +813,7 @@ IDs, as VARCHAR with `stringColumn()`, or as UUID with `uuidColumn()`. See ::: -The standalone `Sender` class exposes only `symbol`, `stringColumn`, +The standalone `Sender` class exposes only these nine: `symbol`, `stringColumn`, `booleanColumn`, `floatColumn`, `intColumn`, `timestampColumn`, `arrayColumn`, `decimalColumn`, and `decimalColumnText`. Its `writer()` method supports every type. @@ -1058,8 +986,10 @@ try { Every sub-array at the same depth must have the same length, and arrays may have 1 to 32 dimensions. Only DOUBLE arrays can be ingested: `longArrayColumn()` exists for protocol parity, but current servers reject it -with `long arrays are not supported, only double arrays`. Query results return -arrays as `{ dimensions, values }`; see +with `long arrays are not supported, only double arrays`. The rejection is +terminal: it stops the sender, and with `sf_dir` the rejected batch blocks the +journal for every table (see [Store-and-forward](#store-and-forward)). Query +results return arrays as `{ dimensions, values }`; see [Reading result values](#reading-result-values). @@ -1212,7 +1142,7 @@ Schema fields: | `geohash(precisionBits)` | GEOHASH | Raw bits, base-32 text of `precisionBits / 5` characters, or `{ bits, precisionBits }` | | `decimal64(scale)`, `decimal128(scale)`, `decimal256(scale)` | DECIMAL | Unscaled `bigint`, decimal text, `number`, or `{ unscaled, scale }` | | `doubleArray()` | DOUBLE[] | Nested `number` arrays, or `{ dimensions, values }` | -| `longArray()` | LONG[] | Encoded for parity; current servers reject LONG arrays | +| `longArray()` | LONG[] | Encoded for parity; current servers reject LONG arrays terminally, as for `longArrayColumn()` | LONG fields take `bigint` so they never lose precision. The object forms (`{ low, high }`, `{ words }`, `{ bits, precisionBits }`, `{ unscaled, scale }`, @@ -1270,6 +1200,19 @@ every later flush fails the same way, and `close()` discards them and rejects with the same error. Call `reset()` to drop every row staged since the last flush, then write the other rows again. +**Batches staged before the first connection.** Until a sender has connected +once, it does not know the server's limit. This applies with +`lazy_connect=on`, with `initial_connect_retry=async`, and to a +store-and-forward restart while QuestDB is down. Batches are then capped only +by `sf_max_segment_bytes`: 4 MiB with `sf_dir`, and no cap without it. A batch +larger than the server's limit passes `flush()` but can never be delivered: +the sender keeps reconnecting, and `waitForAcknowledged()` times out. With +`sf_dir`, the batch also blocks the journal, so later rows are not delivered +and a restarted client fails with `QwpPoolResourceError`. If the client can +start while QuestDB is down and a batch can exceed about 1 MiB, set +`sf_max_segment_bytes=1m`: `2m` is slightly above a default server's limit of +2,097,138 bytes. + ### Closing a sender `close()` publishes the sender's completed rows and discards an unfinished row @@ -1328,7 +1271,10 @@ When a borrowed sender's `close()` fails, the pool discards the sender and opens a new one for the next borrow. Because QuestDB reports rejected batches asynchronously, a sender can fail after its `close()` already succeeded: the error then surfaces on the next borrower's auto-flushing `at()`, `flush()`, or -`close()`, and the pool replaces the sender after that. See +`close()`, and the pool replaces the sender after that. The next borrower's own +staged rows are lost with the failed sender, even rows for other tables: write +them again on a new borrow, as described in +[General usage pattern](#general-usage-pattern). See [Ingestion errors](#ingestion-errors). ### Awaiting acknowledgements @@ -1380,8 +1326,8 @@ try { | Member | Returns | |---|---| | `publishedSequence` | The highest sequence this sender published, including by auto-flushes, or `-1n`. After `flush()`, it covers every row written so far. | -| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout (15 seconds by default), without closing the sender, or with the server's rejection. | -| `acknowledgedSequence` | The highest acknowledged sequence, or `-1n`. | +| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout (15 seconds by default), without closing the sender, or with the server's rejection. It only reads the watermark, so another task can wait while the sender keeps producing. | +| `acknowledgedSequence` | The highest acknowledged sequence, or `-1n`. It never passes a batch that QuestDB rejected. | | `flushAndGetSequence()` | Publishes staged rows and resolves with the highest sequence (`bigint`) this call published, or `-1n` when there was nothing to publish. Rows an earlier auto-flush published are not covered. | :::caution Do not wait on the result of `flushAndGetSequence()` @@ -1411,6 +1357,77 @@ the process exits before the acknowledgement, rows still in memory may be lost; use [store-and-forward](#store-and-forward) to keep them across restarts. +#### Committing source offsets + +To commit offsets in a source such as Kafka only after QuestDB has accepted +the rows, record the sequence of each flushed batch together with the batch's +last source offset. After each flush, commit the newest offset whose sequence +is at or below `acknowledgedSequence`. The watermark is cumulative and never +passes a rejected batch, so a rejection stops further commits, and the next +`flush()` or auto-flushing `at()` reports it. The produce loop never waits for +an individual batch: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// Stand-ins for a source such as a Kafka consumer. +async function* readSource() { + for (let offset = 0n; offset < 2_500n; offset++) { + yield { offset, price: 2615.54, amount: 0.01, timestampMs: Date.now() }; + } +} +async function commitOffset(offset: bigint) { + console.log(`committed through offset ${offset}`); +} + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + // Flushed batches whose last offset waits for QuestDB's acknowledgement. + const pending: { sequence: bigint; offset: bigint }[] = []; + const commitAcknowledged = async () => { + let offset: bigint | undefined; + while (pending.length > 0 && pending[0].sequence <= sender.acknowledgedSequence) { + offset = pending.shift()!.offset; + } + if (offset !== undefined) await commitOffset(offset); + }; + try { + let staged = 0; + let lastOffset = -1n; + const checkpoint = async () => { + await sender.flush(); + pending.push({ sequence: sender.publishedSequence, offset: lastOffset }); + staged = 0; + await commitAcknowledged(); + }; + for await (const event of readSource()) { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", event.price) + .doubleColumn("amount", event.amount) + .at(event.timestampMs, "ms"); + lastOffset = event.offset; + if (++staged === 1_000) await checkpoint(); + } + if (staged > 0) await checkpoint(); + // Before shutting down, wait for the rest and commit it. + await sender.waitForAcknowledged(sender.publishedSequence, 30_000); + await commitAcknowledged(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +To commit as acknowledgements arrive instead of after each flush, register +`ingressSession.onProgress`: it receives events whose `kind` is `published`, +`acknowledged`, or `durable-acknowledged`, with the `sequence` they cover. + ### Transactions By default QuestDB commits each batch on its own. With transactions on, @@ -1475,9 +1492,13 @@ exits. Setting `sf_dir` turns on a disk journal instead: every batch is appended to the journal before it is sent, a background drainer sends it in order, and acknowledged segments are deleted. -Before ingesting, create a deduplicated table while QuestDB is reachable. -Use both the event timestamp and a stable, source-assigned trade ID as upsert -keys: distinct trades can share a millisecond timestamp, symbol, and side. +Before ingesting, create a deduplicated table while QuestDB is reachable, for +example with `lease.query()` as in +[DDL and DML statements](#ddl-and-dml-statements). If the table does not exist +when the first batch arrives, QWP creates it without deduplication, and +replayed batches can then insert duplicate rows. Use both the event timestamp +and a stable, source-assigned trade ID as upsert keys: distinct trades can +share a millisecond timestamp, symbol, and side. Store the trade ID as VARCHAR, not SYMBOL: every trade has its own ID, and SYMBOL is for [bounded sets of values](#column-methods). @@ -1562,19 +1583,23 @@ over the directory's lock (see Lock recovery below). allowances, publishing waits up to `sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, then rejects with `QwpReplayStoreAppendTimeoutError`. -- **Startup.** `lazy_connect=on` lets the pooled client start while QuestDB is - down, as in the example above. `initial_connect_retry=async` alone is not - enough for the pooled client, because its query pool still connects at - startup. A standalone `Sender` needs only `initial_connect_retry=async`. With - the default `off`, the first connection must succeed. +- **Startup.** To start the pooled client while QuestDB is down, see + [Starting while QuestDB is down](#starting-while-questdb-is-down), including + its caveats about locked journals and batch sizes. A standalone `Sender` + needs only `initial_connect_retry=async`. With the default `off`, the first + connection must succeed. - **Lock recovery.** The Node.js client locks a journal directory with a `.lock.owner` directory inside it, which records the owner's host name and process ID, instead of an operating-system file lock. After a crash, a new sender takes over automatically only when the owner ran on the same host and - its process ID is no longer in use. Otherwise the new sender fails with - `QwpReplayStoreLockedError`. This is common in containers: the application - usually runs as process ID 1, which is in use again after a restart, and a - replacement container usually has a different host name. Once you have + its process ID is no longer in use. Otherwise opening the journal fails with + `QwpReplayStoreLockedError`. For the pooled client, `connectQwpNodeClient()` + rejects with it as the `cause` of a `QwpPoolResourceError`, even with + `lazy_connect=on`, so the whole client fails to start, queries included. A + standalone `Sender` rejects on `connect()`. This is common in containers: + the application usually runs as process ID 1, which is in use again after a + restart, and a replacement container usually has a different host name. Once + you have verified that the previous owner has exited and no process is using the slot, remove its stale `//.lock.owner` directory and restart the sender. Here `` is ``, or `-` for a pooled sender. If @@ -1582,12 +1607,23 @@ over the directory's lock (see Lock recovery below). `/.slot-locks/.lock.owner`: this short-lived guard can survive a crash during lock acquisition or quarantine. Remove that specific owner directory only after verifying its owner has exited. Never delete the shared - `.slot-locks` directory or another slot's locks. + `.slot-locks` directory or another slot's locks. Automate this cleanup only + where the deployment guarantees that the previous owner has exited before a + new one starts, for example a single replica that uses the `Recreate` update + strategy and a `ReadWriteOnce` volume. A startup step can then remove the + stale owner directories of the client's own slots before it creates the + client. Anywhere two processes can overlap, recover manually. - **Rejected batches.** A batch that QuestDB rejects terminally, such as one - with a value of the wrong type for an existing column, stays in the journal. - Every new sender on that directory, including the pool's replacement for a - failed pooled sender, sends it again and fails the same way. See - [Ingestion errors](#ingestion-errors) for recovery. + with a value of the wrong type for an existing column, stays at the head of + the journal. Every new sender on that directory, including the pool's + replacement for a failed pooled sender and the same client after a restart, + sends it again and fails the same way, with `QwpReplayRejectedError`, or + `QwpReplayStoreError` while the failed journal closes. Ingestion through the + client therefore stops for every table, not only the table in the rejected + batch. Treat a terminal rejection as an outage: alert on it from + `onSenderError`, then fix the cause or move the journal aside, as described + in [Ingestion errors](#ingestion-errors). `longArrayColumn()` triggers this + on every current server, which rejects LONG arrays terminally. - **Orphans.** With `drain_orphans=on`, a sender also adopts and drains journals with other `sender_id` values left under the same `sf_dir` by processes that crashed, up to `max_background_drainers` (4) at a time. @@ -1741,7 +1777,7 @@ Iterating it with `for await` yields `QwpResultBatch` objects, and the handle's |---|---|---| | `binds` | none | Callback that sets the `$1`, `$2`, ... parameters. See [Bind parameters](#bind-parameters). | | `timeoutMs` | session `queryTimeoutMs` (none) | Deadline that cancels the query. It covers the whole query, including a re-execution after failover. `0` disables it. | -| `initialCredit` | session value (`0`, unbounded) | Flow-control window in bytes. See [Flow control](#flow-control). | +| `initialCredit` | session value (`0`, unbounded) | Flow-control window in bytes. Without one, the client holds whatever the server sends ahead of your loop in memory, so set it for large results. See [Flow control](#flow-control). | | `autoCredit` | `true` | Replenish the credit window as batches are consumed. | | `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | @@ -1980,6 +2016,80 @@ any rows. DDL reports `0`, except `TRUNCATE`, which currently reports Statements run in order on one lease, because each is awaited before the next starts, so a `CREATE TABLE` is complete before the `INSERT` that follows it. +### Read-after-write + +When `flush()` resolves, the client has published the rows, but QuestDB may not +have received them yet. QuestDB acknowledges a batch once it has committed it +to its write-ahead log, and applies committed rows to the table +asynchronously. A query that runs right after ingestion can therefore fail +with `table does not exist` on a first run, or succeed and return no rows. + +When your code must read its own writes, create the table first, write an event +with a unique ID, and poll for **that ID**. Pre-creating the table avoids +mistaking an unrelated SQL error for the first-write table-creation delay. Give +each query the time remaining until the deadline so a stalled query cannot +leave the poll running indefinitely: + +```typescript +import { randomUUID } from "node:crypto"; +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const ddl = await lease.query( + "CREATE TABLE IF NOT EXISTS trades_readback (" + + "timestamp TIMESTAMP, trade_id VARCHAR, symbol SYMBOL" + + ") TIMESTAMP(timestamp) PARTITION BY DAY", + ); + await ddl.completion; + + const tradeId = randomUUID(); + const sender = await db.borrowSender(); + try { + await sender + .table("trades_readback") + .stringColumn("trade_id", tradeId) + .symbol("symbol", "ETH-USD") + .at(Date.now(), "ms"); + await sender.flush(); + } finally { + await sender.close(); + } + + const deadline = Date.now() + 10_000; + let visible = false; + while (!visible) { + const remainingMs = deadline - Date.now(); + if (remainingMs <= 0) throw new Error("trade not visible in time"); + const query = await lease.query( + "SELECT trade_id FROM trades_readback WHERE trade_id = $1 LIMIT 1", + { + binds: (binds) => binds.setVarchar(0, tradeId), + timeoutMs: remainingMs, + }, + ); + for await (const batch of query) visible ||= batch.rowCount > 0; + await query.completion; + if (!visible) { + await new Promise((resolve) => + setTimeout(resolve, Math.min(100, Math.max(0, deadline - Date.now()))), + ); + } + } + console.log(`visible trade: ${tradeId}`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +SQL errors now surface instead of being retried. Do not replace the poll with a +fixed sleep: the apply latency varies with load. + ### Cancellation and timeouts A query ends early in four ways: @@ -2055,10 +2165,13 @@ again, so returning the lease can take up to twice `query_close_timeout_ms`. ### Flow control -By default QuestDB streams results as fast as the network allows, and the -client decodes up to four batches ahead of your loop (`buffer_pool_size`). To -bound how much the server sends ahead, set a byte-credit window with -`initial_credit` in the connect string or `initialCredit` per query: +By default QuestDB streams results as fast as the network allows. The client +decodes up to four batches ahead of your loop (`buffer_pool_size`), but it keeps +every frame it receives in memory until your loop consumes it, so a slow loop +over a large result can hold most of that result in memory. To bound how much +the server sends ahead, set a byte-credit window with `initial_credit` in the +connect string or `initialCredit` per query. Start with 1 MiB for large or +unbounded results: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -2163,6 +2276,24 @@ failover. ## Error handling +Each error leaves the client in a known state. The sections after this table +have the details and examples: + +| Error | Surfaces from | State afterwards | What to do | +|---|---|---|---| +| `TypeError`, `RangeError`, or `Error` from local validation | The column method or `at()` that staged the value | The row in progress is discarded; the sender stays usable | Fix the value and write the row again | +| `QwpBatchTooLargeError` | `flush()`, or the `at()` whose auto-flush sends the batch | The staged rows are kept, and every later flush fails the same way | Call `reset()`, then write the other rows again; see [Flushing](#flushing) | +| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, or an auto-flushing `at()` | The batch stays staged; the sender stays usable | Slow the producer and flush again later; see [Flushing](#flushing) | +| Retriable server rejection | `onSenderError` | The client resends the batch; repeated rejections become terminal | Monitor; no action needed per rejection | +| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls, and with `sf_dir` also `QwpReplayStoreError` | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Ingestion errors](#ingestion-errors) | +| `QwpReconnectExhaustedError` on a sender | `onError` with `terminal: true`, then the next `flush()`, `at()`, or `close()` | The sender has failed; unsent rows are lost | Borrow a new sender; see [Ingestion reconnect](#ingestion-reconnect) | +| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See Lock recovery under [Store-and-forward](#store-and-forward) | +| `QwpPoolResourceError` with another `cause` | `connectQwpNodeClient()`, `borrowSender()`, or `borrowQuery()` | No connection was opened | Unwrap `cause`; see [Connection-level errors](#connection-level-errors) | +| `QwpEgressQueryError` | Query iteration and `completion` | The lease stays usable | Fix the SQL or the bind values | +| `QwpEgressQueryTimeoutError`, `QwpEgressQueryAbandonedError` | Query iteration and `completion` | The lease is busy until QuestDB confirms the cancellation | Close the lease and borrow a new one | +| `QwpEgressQueryCancelTimeoutError` | `completion` | The connection is closed | Close the lease and borrow a new one | +| `QwpReconnectExhaustedError` on a query | Query iteration and `completion` | The lease stays failed, even after QuestDB recovers | Close the lease and borrow a new one; see [Query failover](#query-failover) | + ### Ingestion errors Ingestion reports errors in two ways: @@ -2271,11 +2402,15 @@ automatically after the `close()` that reports the error. What happens to the rejected batch depends on the mode: - **Without store-and-forward**, the failed sender's unacknowledged batches, - including the rejected one, are discarded with it, and the new sender starts + including the rejected one, are discarded with it, and so are rows that a + later borrower staged on it before the error surfaced. The new sender starts empty. -- **With store-and-forward**, the rejected batch stays in the journal. Every - new sender on that directory, including the pool's replacement sender, sends - it again and fails the same way. Fix the cause so that QuestDB accepts the +- **With store-and-forward**, the rejected batch stays at the head of the + journal. Every new sender on that directory, including the pool's + replacement sender and the same client after a restart, sends it again and + fails the same way. Pooled borrows keep getting that journal, so ingestion + through the client stops for every table, not only the table in the + rejected batch, until you act. Fix the cause so that QuestDB accepts the batch, for example by adjusting the table schema, or stop the process and move the journal directory aside. Moving it aside discards every unacknowledged batch in it, not only the rejected one. @@ -2375,8 +2510,10 @@ The pooled client wraps every failure to open a connection, from checking for a specific error. When `addr` lists several hosts, the cause is a `QwpFailoverError` whose `attempts` hold the error of each endpoint. When initial-connect retry is on, for example with a `failover_*` or `reconnect_*` -key or a typed `reconnect` object, the cause is a `QwpReconnectExhaustedError` -instead, and its own `cause` holds the last attempt's error: +key, or with a typed `egressSession.reconnect` object for query connections +(see [Typed reconnect policy](#typed-reconnect-policy)), the cause is a +`QwpReconnectExhaustedError` instead, and its own `cause` holds the last +attempt's error: ```typescript import { @@ -2465,18 +2602,17 @@ queries. Ingestion always needs the primary: replicas refuse writes, and the sender walks the list until it finds the current primary. Queries can use any node. `target` selects which roles queries accept: `any` (the default), `primary`, or -`replica`. Set it with the typed `egress` option, as below: in the connect -string, `target` also filters ingestion (see the caution that follows). It is a -strict filter, not a preference: with `replica`, queries never fall back to the -primary, and they fail when no replica is reachable, including against a single -open source server. Because the pooled client opens a query connection at -startup, `connectQwpNodeClient()` then rejects too, with a -`QwpPoolResourceError` whose `cause` is a `QwpRoleMismatchError`, or a -`QwpFailoverError` holding one per endpoint when `addr` lists several hosts. -With query failover retrying the initial connection, that error is wrapped in a -`QwpReconnectExhaustedError`. To start without a replica, also set -`query_pool_min=0`; queries borrowed before a replica is reachable then reject -with `QwpPoolResourceError`: +`replica`. Set it with the typed `egress` option, as below, because in the +connect string `target` also filters ingestion (see the caution that follows). + +`target` is a strict filter, not a preference. With `replica`, queries never +fall back to the primary, and they fail when no replica is reachable, including +against a single open source server. Because the pooled client opens a query +connection at startup, `connectQwpNodeClient()` then fails too, with a +`QwpPoolResourceError` whose `cause` leads to a `QwpRoleMismatchError`; see +[Connection-level errors](#connection-level-errors) to unwrap it. To start +without a replica, also set `query_pool_min=0`. Queries borrowed before a +replica is reachable then reject with `QwpPoolResourceError`: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -2553,12 +2689,23 @@ endpoint when there is one, and runs the query again from the start: | `failover_backoff_max_ms` | `1000` | Longest delay between retries. | | `failover_max_duration_ms` | `30000` | Time budget per failure. | -When the budget runs out, the query rejects with `QwpReconnectExhaustedError`. -A `QwpEgressQueryError` from the server is a query result and never triggers -failover. Replaying an in-flight `query()` also re-executes DDL and DML: an -`INSERT` may run twice if its completion was lost. For non-idempotent SQL, -use a separate client configured with `failover=off` and check an uncertain -outcome before retrying; see [DDL and DML statements](#ddl-and-dml-statements). +The attempt limit and the time budget apply together, and whichever is reached +first ends the failover. When attempts fail fast, for example with connection +refused while a server restarts, the 8 attempts and their backoff of 50 ms to +1 second, with jitter, take only about 2 to 4 seconds, long before the +30-second budget. To ride out a longer restart, raise `failover_max_attempts`, +or set `maxAttempts: 0` in a typed `egressSession.reconnect` object to remove +the attempt limit and rely on the time budget alone; see +[Typed reconnect policy](#typed-reconnect-policy). + +When failover gives up, the query rejects with `QwpReconnectExhaustedError`, +and the lease stays failed even after QuestDB recovers: close it and borrow a +new one. A `QwpEgressQueryError` from the server is a query result and never +triggers failover. Replaying an in-flight `query()` also re-executes DDL and +DML: an `INSERT` may run twice if its completion was lost. For non-idempotent +SQL, use a separate client configured with `failover=off` and check an +uncertain outcome before retrying; see +[DDL and DML statements](#ddl-and-dml-statements). :::warning Clear partial results when a query restarts @@ -2611,6 +2758,54 @@ request IDs are numbered per connection and every lease of a pooled client shares the callback, so the event cannot tell concurrent queries apart. Use it for logging, and the sequence check above to reset results. +### Typed reconnect policy + +Reconnect and failover behavior comes from the connect-string keys above, or +from typed `reconnect` objects in the second argument of +`connectQwpNodeClient()`: `ingressSession.reconnect` for senders and +`egressSession.reconnect` for queries. You need the object to register +`onEvent` for [connection events](#connection-events). + +:::caution A typed `reconnect` object replaces the connect-string keys + +`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces +the whole reconnect policy parsed from the `reconnect_*`, +`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and +`egressSession.reconnect` replaces the policy parsed from `failover*` keys. +Fields you leave out of the object take the built-in defaults, not the values +from the connect string. When you supply the object, for example to register +`onEvent`, set every bound you rely on in it. + +::: + +The object's fields and the connect-string keys they replace: + +| Field | Ingestion key, default | Query key, default | +|---|---|---| +| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8`. The key accepts `1` or more; the typed field also accepts `0`, unlimited | +| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | +| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | +| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | +| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | +| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | +| `onEvent` | None | None | + +`egressSession: { reconnect: false }` is the typed equivalent of +`failover=off`. + +The two directions treat the first connection differently: + +- **Senders**: setting any `reconnect_*` key makes the first connection retry + within the budget, as if `initial_connect_retry=on`. A typed + `ingressSession.reconnect` object does not, so the first connection still + fails fast. Set `initial_connect_retry` in the connect string to choose the + startup behavior. +- **Query connections**: the first connection retries within the failover + budget, for retryable errors, when you supply an `egressSession.reconnect` + object, set `failover=on` explicitly, or set a `failover_*` key without + `failover=off`. Otherwise it is attempted once. `failover=off` and + `egressSession.reconnect: false` turn reconnects off entirely. + ### Connection events Register `reconnect.onEvent` to observe connections. Events are delivered @@ -2653,13 +2848,8 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { await db.close(); ``` -Supplying an `egressSession.reconnect` options object also makes opening a -query connection retry within the failover budget when the error is retryable. -So does -explicitly setting `failover=on`, or setting a `failover_*` tuning key without -`failover=off`. By contrast, `failover=off` or `egressSession.reconnect: false` -disables the reconnect wrapper. As with other session options, an explicit -`egressSession.reconnect` value replaces the connect-string policy. +Supplying the `reconnect` objects also changes how the first connection is +retried; see [Typed reconnect policy](#typed-reconnect-policy). | Kind | Meaning | |---|---| @@ -2829,8 +3019,10 @@ code: ## ILP transports (legacy) The Node.js `Sender` still ingests over ILP, for existing deployments and for -servers without QWP. ILP senders support HTTP (`http::`, `https::`) and TCP -(`tcp::`, `tcps::`) transports: +servers without QWP. To move ILP code to QWP, see +[From ILP to QWP](#from-ilp-to-qwp); for behavior changes in 5.0.0, see +[Upgrading from 4.x](#upgrading-from-4x). ILP senders support HTTP (`http::`, +`https::`) and TCP (`tcp::`, `tcps::`) transports: ```typescript import { Sender } from "@questdb/nodejs-client"; @@ -2877,7 +3069,8 @@ and the [ILP overview](/docs/connect/compatibility/ilp/overview/). A production-oriented pattern that ingests trades and queries recent prices, with TLS, a token, several hosts, error handling, and failover handling. Before running it, create the deduplicated table on the primary (or reuse the table -from [Store-and-forward](#store-and-forward)): +from [Store-and-forward](#store-and-forward)). If the table is missing, QWP +creates it without deduplication: ```questdb-sql CREATE TABLE IF NOT EXISTS trades_sf ( @@ -2902,6 +3095,7 @@ import { QwpIngressAckTimeoutError, QwpPoolResourceError, QWP_RECONNECT_EVENT_KIND, + QWP_SENDER_ERROR_POLICY, type QwpReconnectEvent, type QwpSenderError, } from "@questdb/nodejs-client"; @@ -2909,6 +3103,11 @@ import { const token = process.env.QDB_TOKEN; if (!token) throw new Error("QDB_TOKEN is not set"); +// Replace with your alerting. +function alertOperator(message: string) { + console.error("ALERT:", message); +} + function logConnection(event: QwpReconnectEvent) { if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { console.info("questdb connection:", event.kind, String(event.endpoint ?? "")); @@ -2928,9 +3127,15 @@ const db = await connectQwpNodeClient( // follows the primary. egress: { target: "replica", compression: "zstd" }, ingressSession: { - onSenderError: (error: QwpSenderError) => - console.error("batch rejected:", error.category, error.serverMessage), - // Terminal failures, such as an exhausted reconnect budget. + onSenderError: (error: QwpSenderError) => { + console.error("batch rejected:", error.category, error.serverMessage); + // A terminally rejected batch stays in the journal and blocks + // ingestion through this client, for every table, until it is fixed. + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + alertOperator(`QuestDB rejected a batch: ${error.serverMessage}`); + } + }, + // Terminal failures, such as a batch that QuestDB rejects terminally. onError: (event) => { if (event.terminal) console.error("ingestion stopped:", event.error); }, @@ -2940,7 +3145,8 @@ const db = await connectQwpNodeClient( egressSession: { queryTimeoutMs: 30_000, // Replaces any failover* keys; omitted fields use the defaults. - reconnect: { maxDurationMs: 30_000, onEvent: logConnection }, + // maxAttempts 0 removes the 8-attempt limit, so failover lasts 30 s. + reconnect: { maxAttempts: 0, maxDurationMs: 30_000, onEvent: logConnection }, onReplayReset: (event) => console.warn("query restarts on", String(event.endpoint)), }, From 1d13aecef8e4c88c0f0779f76b251d9a9fd5bc5f Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 11:50:34 +0100 Subject: [PATCH 14/25] docs(clients): document the current sender error policies for Java and .NET The Java and .NET pages still described DROP_AND_CONTINUE and HALT, which contradicted the corrected store-and-forward error table. - Java: list all ten categories and the RETRIABLE, RETRIABLE_OTHER, TERMINAL, and ABANDONED policies, and add getQuarantinedPath() - .NET: document Retriable, Terminal, and Abandoned with their default categories, the accepted on_*_error values (drop now means retry), and that the keys can only make a category stricter - Store-and-forward tuning: log levels follow terminal and retriable policies - Changelog: mention the Java and .NET page updates --- documentation/changelog.mdx | 2 +- documentation/connect/clients/dotnet.md | 57 +++++++++++-------- documentation/connect/clients/java.md | 12 +++- .../store-and-forward/operating-and-tuning.md | 5 +- 4 files changed, 46 insertions(+), 30 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index ade7dc3dc9..388b61bc06 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -31,7 +31,7 @@ This page tracks significant updates to the QuestDB documentation. ### Updated - [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ -- [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and [client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection +- [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and the [Java](/docs/connect/clients/java/#ingestion-errors) and [.NET](/docs/connect/clients/dotnet/#how-errors-surface) client pages now document the same policies. [Client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/connect/clients/dotnet.md b/documentation/connect/clients/dotnet.md index 684ea97cd3..5f4d5318fa 100644 --- a/documentation/connect/clients/dotnet.md +++ b/documentation/connect/clients/dotnet.md @@ -783,10 +783,12 @@ Each error is classified into a `SenderErrorCategory` and assigned a | Policy | Effect | Default categories | |---|---|---| -| `DropAndContinue` | The rejected batch is dropped; the sender keeps running. | `SchemaMismatch`, `WriteError` | -| `Halt` | The sender latches terminal; the next producer call throws `LineSenderServerException`. | `ParseError`, `InternalError`, `SecurityError`, `ProtocolViolation`, `Unknown` | +| `Retriable` | The client reconnects and resends the batch; the sender keeps running. A batch that keeps being rejected escalates to `Terminal`. | `WriteError`, `InternalError`, `NotWritable`, `DictionaryGap`, `Unknown` | +| `Terminal` | The sender latches terminal; the next producer call throws `LineSenderServerException`. | `SchemaMismatch`, `ParseError`, `SecurityError`, `ProtocolViolation` | +| `Abandoned` | Store-and-forward data that can never be sent was set aside at `QuarantinedPath` and is not retried. | `DataLoss` | -After a `Halt`, discard the sender and create a new one. +After a `Terminal` error, discard the sender and create a new one. There is no +drop policy: the client never discards a rejected batch silently. ### Error handler @@ -814,8 +816,8 @@ Each `SenderError` carries the following fields: | Field | Description | |---|---| -| `Category` | `SchemaMismatch`, `ParseError`, `InternalError`, `SecurityError`, `WriteError`, `ProtocolViolation`, `Unknown`. Use for programmatic dispatch. | -| `AppliedPolicy` | `DropAndContinue` (batch dropped, sender continues) or `Halt` (sender latched terminal; next API call throws `LineSenderServerException`). | +| `Category` | `SchemaMismatch`, `ParseError`, `InternalError`, `SecurityError`, `WriteError`, `NotWritable`, `DictionaryGap`, `ProtocolViolation`, `DataLoss`, `Unknown`. Use for programmatic dispatch. | +| `AppliedPolicy` | `Retriable` (batch resent, sender continues), `Terminal` (sender latched terminal; next API call throws `LineSenderServerException`), or `Abandoned` (only with `DataLoss`: the data was set aside). | | `ServerStatusByte` | Raw QWP status byte (e.g. `0x03` for `SchemaMismatch`). `-1` (`SenderError.NoStatusByte`) on `ProtocolViolation` and engine-internal terminal failures. | | `ServerMessage` | Human-readable server text (≤ 1024 UTF-8 bytes), or `null`. See [Message stability](#message-stability) and [PII safety](#message-pii). | | `MessageSequence` | Server's per-frame QWP wire sequence for the error frame. `-1` (`SenderError.NoMessageSequence`) for engine-internal failures. **Resets on reconnect** — only meaningful within one connection. | @@ -824,6 +826,7 @@ Each `SenderError` carries the following fields: | `DetectedAtUtc` | Wall-clock receipt time on the I/O thread; for ops timelines, not for correlation. | | `Exception` | Non-`null` for engine-internal failures (connect-budget exhaustion, fatal upgrade reject); `null` for server rejections. | | `IsInitialConnect` | `true` if the engine never reached a first successful connection (config / connectivity issue); always `false` for server-side rejections. | +| `QuarantinedPath` | For `DataLoss`, where the set-aside data was preserved; `null` otherwise. | #### Message stability {#message-stability} @@ -857,8 +860,8 @@ and the `(MessageSequence, FromFsn, ToFsn)` triple. ### Synchronous errors Misconfiguration and API-misuse errors surface synchronously as `IngressError` -(or its subclass `LineSenderServerException` for HALT-policy server -rejections). They are thrown directly from the call site: +(or its subclass `LineSenderServerException` for server rejections with the +`Terminal` policy). They are thrown directly from the call site: | Site | Throws when | |---|---| @@ -869,7 +872,7 @@ rejections). They are thrown directly from the call site: | Array `Column(...)` overloads | The `shape` does not match the element count, dimensionality exceeds 32, or the element type is not `double` / `long`. | | `ColumnGeohash(...)` | `precisionBits` is outside `[1, 60]`. | | `ColumnDecimal*(...)` with explicit `scale` | `scale` is outside `[0, 18]` (DECIMAL64), `[0, 38]` (DECIMAL128), or `[0, 76]` (DECIMAL256). | -| Producer-thread call after `Halt` policy fired | The next `Table`, `Column`, `AtAsync`, or `SendAsync` throws `LineSenderServerException` carrying the latched `SenderError`. Discard the sender and create a new one. | +| Producer-thread call after a `Terminal` policy fired | The next `Table`, `Column`, `AtAsync`, or `SendAsync` throws `LineSenderServerException` carrying the latched `SenderError`. Discard the sender and create a new one. | Authentication failures surface differently between paths: a `401` / `403` during the WebSocket upgrade returns synchronously from `Sender.New` / @@ -880,26 +883,31 @@ the sender latched terminal. ### Per-category policy -Override the default policy per category with the `on_*_error` connect-string -keys (values `halt` or `drop`): +Make a category stricter with the `on_*_error` connect-string keys. The +accepted values are `halt` (alias `terminal`) and `retry` (alias `retriable`). +The legacy values `drop` and `drop_and_continue` now mean `retry`, because the +client never drops a batch. Any other value, such as `auto` or +`retriable_other`, is rejected with a `ConfigError`: ```csharp -// Treat a schema mismatch as fatal instead of dropping the batch. +// Stop the sender on write errors instead of retrying them. using var sender = Sender.New( - "ws::addr=localhost:9000;on_schema_mismatch_error=halt;"); + "ws::addr=localhost:9000;on_write_error=halt;"); ``` | Key | Scope | |---|---| -| `on_server_error` | Catch-all default for every category. | -| `on_schema_mismatch_error` (alias: `on_schema_error`) | Schema-validation rejections. | -| `on_parse_error` | Client-side parse errors. | -| `on_internal_error` | Unexpected client-side errors. | -| `on_security_error` | Auth / TLS errors. | -| `on_write_error` | Transport write failures. | - -`ProtocolViolation` and `Unknown` are always `Halt`, regardless of these keys. -For programmatic control, set `SenderOptions.error_policy_resolver` to a +| `on_server_error` | Catch-all default for every category below. | +| `on_schema_mismatch_error` (alias: `on_schema_error`) | `SchemaMismatch`: the batch does not match the table schema. | +| `on_parse_error` | `ParseError`: the server could not parse the batch. | +| `on_internal_error` | `InternalError`: an unexpected server-side failure. | +| `on_security_error` | `SecurityError`: the server denied the write. | +| `on_write_error` | `WriteError`: the write failed, for example because the table is not accepting writes. | + +The keys can only make a category stricter. `SchemaMismatch`, `ParseError`, +`SecurityError`, and `ProtocolViolation` are always `Terminal`, even if their +key says `retry`, and `Unknown` stays `Retriable` regardless of these keys. For +programmatic control, set `SenderOptions.error_policy_resolver` to a `SenderErrorPolicyResolver` delegate. ### Connection-level errors @@ -934,9 +942,10 @@ A summary of how the engine treats each error class on the wire: | Auth (`401` / `403`) on any endpoint | Terminal | Halts the failover loop immediately; the sender / query client latches non-recoverable. | | Role reject (`421` + `X-QuestDB-Role`) | Topology-level (transient if `PRIMARY_CATCHUP`, otherwise terminal for the loop) | The client tries the next endpoint; if every endpoint rejects, surfaces as `QwpRoleMismatchException` (egress) or the sender's reconnect loop exhausts. | | Version mismatch during upgrade | Per-endpoint, **not** terminal | The client moves on to the next endpoint. | -| Server rejection of a batch (`SchemaMismatch`, `ParseError`, `WriteError`, etc.) | Per the `on_*_error` policy — default is `DropAndContinue` for `SchemaMismatch` / `WriteError`, `Halt` for everything else. | `DropAndContinue` keeps the sender alive; `Halt` latches the sender so the next producer call throws `LineSenderServerException`. | +| Server rejection of a batch (`SchemaMismatch`, `ParseError`, `WriteError`, etc.) | Per category: `Retriable` for `WriteError`, `InternalError`, `NotWritable`, and `DictionaryGap`; `Terminal` for `SchemaMismatch`, `ParseError`, and `SecurityError`. The `on_*_error` keys can make a retriable category terminal. | `Retriable` resends the batch and keeps the sender alive; `Terminal` latches the sender so the next producer call throws `LineSenderServerException`. | | TCP / TLS failure, `404`, `503`, mid-stream drop | Transient | Fed into the ingress reconnect loop (`reconnect_max_*` keys) or, on egress, the per-query failover loop (`failover_*` keys). | -| `ProtocolViolation`, `Unknown` | Terminal | Always `Halt`, regardless of `on_*_error` settings. | +| `ProtocolViolation` | Terminal | Always `Terminal`, regardless of `on_*_error` settings. | +| `Unknown` (a status this client does not know) | Retriable | Resent rather than stopping the sender, regardless of `on_*_error` settings. | ### Connection events @@ -962,7 +971,7 @@ Event kinds: `Connected`, `Disconnected`, `Reconnected`, `FailedOver`, `AuthFailed` and `ReconnectBudgetExhausted` are **terminal**: the sender latches a non-recoverable failure, the next producer-thread call (`Table`, `Column`, `AtAsync`, `SendAsync`) throws `IngressError` (or -`LineSenderServerException` if a HALT-policy error was latched alongside), +`LineSenderServerException` if a `Terminal`-policy error was latched alongside), and no further data can be sent. Discard the sender, build a new one, and replay any state your application owns. `DroppedConnectionNotifications` on `IQwpWebSocketSender` counts events that were dropped because a slow listener diff --git a/documentation/connect/clients/java.md b/documentation/connect/clients/java.md index 1df7d8239f..020dbdd590 100644 --- a/documentation/connect/clients/java.md +++ b/documentation/connect/clients/java.md @@ -1242,18 +1242,24 @@ regardless of how the sender was created: | Field | Accessor | Description | |-------|----------|-------------| -| Category | `getCategory()` | `SCHEMA_MISMATCH`, `PARSE_ERROR`, `INTERNAL_ERROR`, `SECURITY_ERROR`, `WRITE_ERROR`, `PROTOCOL_VIOLATION`, or `UNKNOWN` | -| Policy | `getAppliedPolicy()` | `DROP_AND_CONTINUE` (batch dropped, sender continues) or `HALT` (next API call throws `LineSenderServerException`) | +| Category | `getCategory()` | `SCHEMA_MISMATCH`, `PARSE_ERROR`, `INTERNAL_ERROR`, `SECURITY_ERROR`, `WRITE_ERROR`, `NOT_WRITABLE`, `DICTIONARY_GAP`, `PROTOCOL_VIOLATION`, `DATA_LOSS`, or `UNKNOWN` | +| Policy | `getAppliedPolicy()` | `RETRIABLE` (the client reconnects and resends the batch), `RETRIABLE_OTHER` (resends it to another endpoint), `TERMINAL` (the sender stops; the next API call throws `LineSenderServerException`), or `ABANDONED` (only with `DATA_LOSS`: store-and-forward data that can never be sent was set aside) | | Server message | `getServerMessage()` | Human-readable error text from the server (may be null) | | Table name | `getTableName()` | The rejected table (null for multi-table batches) | | FSN range | `getFromFsn()` / `getToFsn()` | Frame sequence number span identifying the rejected batch | | Message sequence | `getMessageSequence()` | Server's per-frame sequence number (`-1` if not available) | | Status byte | `getServerStatusByte()` | Raw QWP status code (`-1` if not available) | +| Quarantined path | `getQuarantinedPath()` | For `DATA_LOSS`, where the set-aside data was preserved (null otherwise) | + +There is no drop policy: a rejected batch is resent, stops the sender, or, for +`DATA_LOSS` only, is set aside. See +[Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames) +for the default policy of each category. The error handler runs on a dedicated dispatcher thread, never on the I/O or producer thread. -When a sender owned by `QuestDB` enters a terminal `HALT` state, the next +When a sender owned by `QuestDB` stops with a `TERMINAL` error, the next producer-thread call throws `LineSenderServerException`. The pool detects the failure on close/return and replaces the failed sender with a fresh one on the next borrow. diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index 83fcafd497..f32ab071ca 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -312,8 +312,9 @@ A conformant client exposes at minimum: by category. Background `SCHEMA_MISMATCH` is usually a schema-drift symptom worth alerting on. -The default error handler logs every received `SenderError` — -`ERROR`-level for HALT, `WARN`-level for DROP. Replace it only if you +The default error handler logs every received `SenderError`: +`ERROR`-level for terminal and abandoned errors, `WARN`-level for retriable +ones. Replace it only if you are also routing the errors somewhere else (Sentry, structured logs): silence is forbidden by the contract. From 3289438da8be89cb70840c34791f3b57258864f3 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 14:22:19 +0100 Subject: [PATCH 15/25] docs(nodejs): clarify QWP client usage and delivery semantics --- .github/workflows/validate-links.yml | 3 + documentation/changelog.mdx | 9 +- documentation/concepts/delivery-semantics.md | 19 +- .../connect/clients/connect-string.md | 81 +- .../clients/date-to-timestamp-conversion.md | 12 +- documentation/connect/clients/java.md | 17 +- documentation/connect/clients/nodejs.md | 1134 +++++++++++------ .../connect/compatibility/pgwire/nodejs.md | 5 +- documentation/connect/overview.md | 5 + .../wire-protocols/qwp-client-behavior.md | 36 +- .../wire-protocols/qwp-ingress-websocket.md | 8 +- .../client-failover/concepts.md | 29 +- .../client-failover/configuration.md | 9 +- .../store-and-forward/concepts.md | 56 +- .../store-and-forward/configuration.md | 15 +- .../store-and-forward/operating-and-tuning.md | 43 +- .../store-and-forward/when-to-use.md | 11 +- package.json | 3 +- plugins/raw-markdown/convert-components.js | 6 +- .../raw-markdown/convert-components.test.js | 23 + shared/clients.json | 2 +- 21 files changed, 986 insertions(+), 540 deletions(-) diff --git a/.github/workflows/validate-links.yml b/.github/workflows/validate-links.yml index bb51f12c7f..3cea5447c3 100644 --- a/.github/workflows/validate-links.yml +++ b/.github/workflows/validate-links.yml @@ -36,5 +36,8 @@ jobs: - name: Install dependencies run: yarn install --frozen-lockfile + - name: Run plugin tests + run: yarn test + - name: Build site for broken link validation run: yarn build diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 388b61bc06..3f676f3054 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -8,6 +8,13 @@ description: Recent updates and improvements to the QuestDB documentation. This page tracks significant updates to the QuestDB documentation. +## October 2026 + +### Updated + +- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ +- [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and the [Java](/docs/connect/clients/java/#ingestion-errors) and [.NET](/docs/connect/clients/dotnet/#how-errors-surface) client pages now document their retriable, terminal, and abandoned policies. [Client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection + ## September 2026 ### QuestDB Enterprise Releases @@ -30,8 +37,6 @@ This page tracks significant updates to the QuestDB documentation. ### Updated -- [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ -- [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and the [Java](/docs/connect/clients/java/#ingestion-errors) and [.NET](/docs/connect/clients/dotnet/#how-errors-surface) client pages now document the same policies. [Client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection - [query_activity()](/docs/query/functions/meta/#query_activity) - Documented the `is_wal`, `memory_used`, and `memory_limit` columns - [wal_tables()](/docs/query/functions/meta/#wal_tables) - Documented the `errorTag`, `errorMessage`, and `memoryPressure` columns; `errorTag` reads `OUT OF MEMORY` after a WAL apply memory limit breach - [SHOW](/docs/query/sql/show/) - `SHOW USERS`, `SHOW GROUPS`, and `SHOW SERVICE ACCOUNTS`, including their filtered forms, gain a trailing `memory_limit` column. Clients that read these results by position need [updating](/docs/security/rbac/#memory-limit-upgrade) diff --git a/documentation/concepts/delivery-semantics.md b/documentation/concepts/delivery-semantics.md index 597c1910ca..0f7a791ad2 100644 --- a/documentation/concepts/delivery-semantics.md +++ b/documentation/concepts/delivery-semantics.md @@ -7,10 +7,19 @@ description: exactly-once outcomes. --- -QuestDB clients deliver data **at-least-once**: every row your application -publishes is guaranteed to reach the server, but under failure it may arrive -more than once. Storing each row exactly once is the application's -responsibility, and QuestDB provides the mechanisms to make it routine. +QuestDB clients deliver data **at-least-once**: a sender keeps every row your +application publishes until the server acknowledges it, and resends it after a +failure, so under failure a row may arrive more than once. Storing each row +exactly once is the application's responsibility, and QuestDB provides the +mechanisms to make it routine. + +The guarantee holds while the sender runs. Without +[store-and-forward](/docs/high-availability/store-and-forward/concepts/), +unacknowledged rows live in memory and are lost if the process exits, or the +sender closes, before the server acknowledges them. A Node.js sender in +default memory mode also gives up after `reconnect_max_duration_millis` of +outage; see the +[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). This page explains where duplicates come from and how to suppress them. @@ -19,7 +28,7 @@ This page explains where duplicates come from and how to suppress them. | Property | Meaning | Where it comes from | |----------|---------|---------------------| | **At-most-once** | Each row reaches the server zero or one times. Rows can be lost. | A "fire and forget" client that does not retransmit on failure. | -| **At-least-once** | Each row reaches the server one or more times. No row is lost; duplicates are possible. | A client that retransmits unacknowledged data after a transport error. **This is the QuestDB client default.** | +| **At-least-once** | Each row reaches the server one or more times. No row is lost; duplicates are possible. | A client that retransmits unacknowledged data after a transport error. **QuestDB clients provide this while the sender runs**, and across restarts with store-and-forward. | | **Exactly-once** | Each row is stored exactly once. | At-least-once delivery plus server-side deduplication on a key covering row identity. | QuestDB's clients retransmit unacknowledged batches after transport errors, diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 1f0ab5baff..0e60f12f66 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -20,6 +20,8 @@ configures both without edits. The Node.js client is the exception for `target` and `zone`, which it also applies to ingress; see [Role filter and zone preference](#role-filter-and-zone-preference). The *Applies to:* tag on each section below marks which direction a key affects. +The [Node.js client page](/docs/connect/clients/nodejs/#differences-from-other-clients) +lists every key where that client differs from this reference. For legacy InfluxDB Line Protocol (ILP) transports (`http`, `https`, `tcp`, `tcps`), see the [ILP overview](/docs/connect/compatibility/ilp/overview/). @@ -147,7 +149,7 @@ wss::addr=node-a:443,node-b:443;target=replica;zone=eu-west-1a; ``` Senders in other clients ignore `target`, so they can share this string. On the -Node.js pooled client, `target=replica` in the connect string stops ingestion; +Node.js client, `target=replica` in the connect string stops ingestion; set the role with the typed `egress` option instead, as described under [Role filter and zone preference](#role-filter-and-zone-preference). @@ -180,7 +182,7 @@ caveats), follow the section links from the [Key index](#key-index). | Zone-aware routing with DR last-resort | egress | `zone=` | `target` | | Tune ingest batching | ingress | — | Clients with auto-flush: `auto_flush_rows`, `auto_flush_interval`, `auto_flush_bytes` | | Disable auto-flush (manual `flush()` only) | ingress | `auto_flush=off` | — | -| Memory-buffered ingest (no disk durability) | ingress | (omit `sf_dir`) | `init_buf_size`, `max_buf_size` | +| Memory-buffered ingest (no disk durability) | ingress | (omit `sf_dir`) | Node.js: `sf_max_total_bytes` (memory replay capacity), `auto_flush_rows` (batching); clients with row-buffer sizing: `init_buf_size`, `max_buf_size` | | Durable store-and-forward ingest | ingress | `sf_dir` | `sender_id`, `sf_max_segment_bytes`, `sf_max_total_bytes`, `sf_append_deadline_millis` | | Run multiple senders sharing one `sf_dir` | ingress | `sf_dir`, `sender_id` | unique `sender_id` per sender | | Orphan recovery for crashed senders | ingress | `drain_orphans=on` | `max_background_drainers` | @@ -229,12 +231,14 @@ WebSocket upgrade request. deployments. - `auth_timeout_ms` — per-host upper bound on the upgrade response read. Does not cover TLS handshake or post-upgrade frame reads, which use OS or - hard-coded defaults. Default: `15000` (15 s). + hard-coded defaults. Default: `15000` (15 s). On Node.js, it defaults to + `connect_timeout` when only that key is set. - `connect_timeout` — integer milliseconds, must be `> 0`. Applies to ingress and egress. Bounds the TCP connect phase for each endpoint, so a black-holed host in a multi-host `addr` no longer stalls the [endpoint walk](#failover-keys) until the OS connect timeout. Unset by - default in most clients; Node.js defaults to `15000` (15 seconds). + default in most clients, which then use the OS timeout. Node.js defaults it + to `15000` (15 seconds) and also bounds DNS and the TLS handshake with it. **Mutual TLS (mTLS).** Not supported. The client validates the server's certificate against a trust store but cannot present a client certificate; @@ -516,8 +520,9 @@ equivalent — same architecture, no durability across restarts. - Taken verbatim. Absolute paths recommended for production; relative paths resolve against the process working directory. - The client does **not** expand shell-style syntax such as `~`. - - The client creates the leaf directory if it is missing, but the parent - must already exist — it does not create paths recursively. + - Create `sf_dir` before opening the sender unless you use Node.js, which + creates the slot and any missing parent directories recursively. Other + clients may create only the leaf slot directory. - `sender_id` — slot identity. The slot lives at `//`, used verbatim as the directory name. Allowed characters: letters, digits, `_`, `-`. No path separators, no `.`, no spaces. Two senders @@ -557,7 +562,7 @@ equivalent — same architecture, no durability across restarts. On Node.js with `sf_dir`, this is a journal size target, not a hard disk limit: transaction-closing batches and retained symbol dictionaries can exceed it, and other metadata needs additional space. Provision disk - headroom; see the [Node.js capacity guidance](/docs/connect/clients/nodejs/#store-and-forward). + headroom; see the [Node.js capacity guidance](/docs/connect/clients/nodejs/#sf-capacity). Without `sf_dir`, the key caps the in-memory replay queue. ### Sender restart and replay @@ -592,7 +597,7 @@ The Node.js client does not use an OS lock. It locks a slot with a locked, and a new sender then fails with `QwpReplayStoreLockedError`. Node.js and other clients do not see each other's locks: never let them use the same `sf_dir` at the same time. See the -[Node.js client](/docs/connect/clients/nodejs/#store-and-forward) for lock +[Node.js client](/docs/connect/clients/nodejs/#sf-lock-recovery) for lock recovery. ::: @@ -648,9 +653,8 @@ SF mode and memory-only mode share the same loop. A **running** sender retries a transport outage indefinitely with capped exponential backoff — there is no wall-clock give-up: the whole point of the buffering architecture is that a producer survives an arbitrarily long outage. The -exception is a Node.js sender with neither `sf_dir` nor background replay -(`initial_connect_retry=async`, or pooled `lazy_connect=on`), which gives up -after `reconnect_max_duration_millis`; see below. +exception is a Node.js sender in default memory mode, which gives up after +`reconnect_max_duration_millis`; see below. - `reconnect_initial_backoff_millis` — initial wait between reconnect attempts. Backoff grows exponentially up to `reconnect_max_backoff_millis`. @@ -664,10 +668,10 @@ after `reconnect_max_duration_millis`; see below. constructor gives up and returns the error. The running loop and the `async` initial connect never consult it. Default: `300000` (5 min). Setting this enables `initial_connect_retry=on` implicitly; see below. - The Node.js client differs: a sender with neither `sf_dir` nor background - replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies - this budget to every outage, and fails with `QwpReconnectExhaustedError` - when it runs out. See the + The Node.js client differs: a sender in default memory mode, without + `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies this + budget to every outage, and fails with `QwpReconnectExhaustedError` when it + runs out. See the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - `initial_connect_retry` — whether the client retries the initial connect attempt on failure. @@ -690,7 +694,7 @@ after `reconnect_max_duration_millis`; see below. milliseconds waiting for buffered frames to drain. Set to `0` or `-1` for fast close (skip the drain). **The default differs by client**: `60000` (60 s) on Java and .NET, `5000` (5 s) on Rust, C, C++ and Python, which - share the same Rust core, and on Node.js. + share the same Rust core, and on Go and Node.js. This is the shutdown data-loss window. Setting it to `0` skips the drain entirely and drops un-ACKed batches on every clean shutdown. @@ -698,10 +702,11 @@ after `reconnect_max_duration_millis`; see below. Authentication rejection (HTTP `401` / `403`) never moves the loop to another host. Before a sender's first successful connection it is terminal in every client. After that, clients differ: the Java client retries it indefinitely, -the Node.js client does so for senders with `sf_dir` or background replay -(`initial_connect_retry=async`, or pooled `lazy_connect=on`), and the Rust, C, +the Node.js client does so for senders with `sf_dir` or in background memory +mode (`initial_connect_retry=async` or `lazy_connect=on`), and the Rust, C, C++, Python, Go, and .NET clients stop. Query connections and orphan drainers -always stop. See +stop too, except a Java orphan drainer whose token comes from a token +provider, which retries for a bounded time before quarantining the slot. See [authentication during failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). ### Egress failover {#egress-failover} @@ -765,7 +770,7 @@ connect string without an "unknown configuration key" error — the Sender does not interpret the values. Range, enum, and type checks happen on the egress side; the Sender silently accepts even a value the `QwpQueryClient` parser would reject. The Node.js `Sender` accepts these keys -too, but logs a warning that it ignores them. +too and logs a warning for the ones it ignores; it applies `client_id`. - `compression` — result-batch compression the client advertises. Options: `raw` (default — no compression; the client omits the accept-encoding @@ -811,9 +816,10 @@ per-language names. `QuestDBClient.Connect`, `connectQwpNodeClient`).* Every client now leads with a pooled facade, so these keys are a first-contact -concern. The `Sender` and query-client parsers accept and ignore them (the -Node.js `Sender` logs a warning that it ignores them); the facade reads them -off the string. Each has an equivalent builder setter, and an +concern. The `Sender` and query-client parsers accept and ignore them; the +facade reads them off the string. The Node.js `Sender` logs a warning for the +pool keys it ignores, and applies `lazy_connect`, which starts it in +background memory mode. Each has an equivalent builder setter, and an explicit setter always wins over the string. - `sender_pool_min` — senders kept open even when idle. `0` lets the pool close @@ -852,11 +858,22 @@ consumed by the application. :::caution Accepted, but not applied by every client Every client's parser accepts the six `on_*_error` keys below, but only -clients that implement the policy layer act on them. The Go and .NET clients -apply them. **In the Java reference client and the Node.js, Rust, C, C++, and -Python clients they are currently accepted no-ops**, so setting -`on_write_error=retriable_other` parses cleanly and changes nothing. The -category table and precedence model below describe the target contract. +clients that implement the policy layer act on them, and the two that do +accept different values: + +- **Go** applies them as described here. `on_server_error` accepts `auto`, + `terminal`, `retriable`, or `retriable_other`, and the per-category keys + accept the same values except `auto`. +- **.NET** accepts only `halt` (alias `terminal`) and `retry` (alias + `retriable`), maps the legacy `drop` and `drop_and_continue` to `retry`, and + rejects any other value, including `auto` and `retriable_other`, with a + `ConfigError`. Its keys can only make a retriable category terminal; see the + [.NET client](/docs/connect/clients/dotnet/#per-category-policy). +- **In the Java reference client and the Node.js, Rust, C, C++, and Python + clients they are currently accepted no-ops**, so setting + `on_write_error=retriable_other` parses cleanly and changes nothing. + +The category table and precedence model below describe the target contract. ::: @@ -898,7 +915,9 @@ poison-frame detector (`max_frame_rejections`, default `4`). `PROTOCOL_VIOLATION` is always terminal and `UNKNOWN` always retriable (fail open: a status byte from a newer server degrades to retry, not to a dead -sender); neither can be overridden. Per-client wiring of the override surface +sender); the `on_*_error` keys cannot override either. The .NET client's +programmatic resolver is the exception: it can change the policy for +`UNKNOWN`. Per-client wiring of the override surface may lag the spec — check your client's documentation for which of the resolver / per-category / connect-string layers it exposes. For the full model see the @@ -922,7 +941,7 @@ description and behaviour notes. | `buffer_pool_size` | int (≥ 1) | `4` | [Query client keys](#egress-keys) | | `catch_up_cap_gap_min_escalation_window_millis` | int (ms) | `300000` (5 min) | [Store-and-forward](#sf-keys) | | `client_id` | string | client-specific | [Query client keys](#egress-keys) | -| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, Node.js `5000` | [Ingress reconnect](#reconnect-keys) | +| `close_flush_timeout_millis` | int (ms) | Java/.NET `60000` / Rust, C, C++, Python, Go, Node.js `5000` | [Ingress reconnect](#reconnect-keys) | | `compression` | enum (`raw` / `zstd` / `auto`) | `raw` | [Query client keys](#egress-keys) | | `compression_level` | int (`1`–`22`) | `1` | [Query client keys](#egress-keys) | | `connect_timeout` | int (ms, `> 0`) | unset (Node.js: `15000`) | [Authentication](#auth) | @@ -956,7 +975,7 @@ description and behaviour notes. | `on_write_error` | enum | `retriable` | [Error handling](#error-handling) | | `pass` | string | unset | [Authentication](#auth) (alias of `password`) | | `password` | string | unset | [Authentication](#auth) | -| `poison_min_escalation_window_millis` | int (ms) | `5000` (Node.js: `300000`) | [Error handling](#error-handling) | +| `poison_min_escalation_window_millis` | int (ms) | `5000` (Node.js: `300000`; Go: not supported, its window is `reconnect_max_duration_millis`) | [Error handling](#error-handling) | | `query_close_timeout_ms` | int (ms) | `5000` | [Query client keys](#egress-keys) | | `query_pool_max` | int | `4` | [Connection pool](#pool-keys) | | `query_pool_min` | int | `1` | [Connection pool](#pool-keys) | diff --git a/documentation/connect/clients/date-to-timestamp-conversion.md b/documentation/connect/clients/date-to-timestamp-conversion.md index 4fc51b5d63..dc401b259e 100644 --- a/documentation/connect/clients/date-to-timestamp-conversion.md +++ b/documentation/connect/clients/date-to-timestamp-conversion.md @@ -331,17 +331,17 @@ no arithmetic: pass `getTime()` with the `"ms"` unit. ```javascript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const tradeDate = new Date("2024-08-05T00:00:00Z"); +const settlementDate = new Date("2024-08-05T00:00:00Z"); const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); try { const sender = await db.borrowSender(); try { await sender - .table("trades") + .table("settlements") .symbol("symbol", "ETH-USD") - .timestampColumn("trade_date", tradeDate.getTime(), "ms") - .doubleColumn("price", 2615.54) + .timestampColumn("settlement_date", settlementDate.getTime(), "ms") + .doubleColumn("amount", 0.5) .at(Date.now(), "ms"); } finally { await sender.close(); @@ -352,8 +352,8 @@ try { ``` For an explicit microsecond value, convert through `bigint`: -`BigInt(tradeDate.getTime()) * 1000n`. Nanosecond timestamps, with the `"ns"` -unit, must be a `bigint`. +`BigInt(settlementDate.getTime()) * 1000n`. Nanosecond timestamps, with the +`"ns"` unit, must be a `bigint`. Learn more about the [QuestDB Node.js Client](/docs/connect/clients/nodejs/) diff --git a/documentation/connect/clients/java.md b/documentation/connect/clients/java.md index 020dbdd590..cdee6b7aed 100644 --- a/documentation/connect/clients/java.md +++ b/documentation/connect/clients/java.md @@ -215,7 +215,8 @@ apply latency varies with load. This applies to every client, not just Java. See the equivalent poll in the [Python](/docs/connect/clients/python/), [Rust](/docs/connect/clients/rust/) and -[Go](/docs/connect/clients/go/) quick starts. +[Go](/docs/connect/clients/go/) quick starts, and in the Node.js client's +[Read-after-write](/docs/connect/clients/nodejs/#read-after-write) section. The `QuestDB` handle is a facade over two distinct kinds of client: a [`Sender`](#data-ingestion) for ingestion (`db.borrowSender()`) and a @@ -1325,8 +1326,10 @@ What is and isn't carried on `onError`: ### Connection-level errors - **Authentication failure**: `401`/`403` HTTP response before the WebSocket - upgrade completes. Terminal across all endpoints. The borrow that - triggered the connect rethrows `LineSenderException`. + upgrade completes. Terminal across all endpoints for the query client and + for a sender's first connection: the borrow that triggered the connect + rethrows `LineSenderException`. A sender that has connected once retries + instead; see [Which failures are retried](#which-failures-are-retried). - **Malformed frames**: `QwpDecodeException` or WebSocket close with a terminal code. - **Role mismatch**: `QwpRoleMismatchException` when all endpoints report @@ -1427,8 +1430,12 @@ client and gives you a fresh one on your next query. ### Which failures are retried -- **Authentication failures** (a bad token or wrong credentials) stop - immediately on every host — retrying cannot help. +- **Authentication failures** (a bad token or wrong credentials) stop the + query client, and a sender's first connection, immediately on every host. A + sender that has connected once retries them indefinitely, keeping its data + buffered, and reports each rejection to its error handler as a `RETRIABLE` + `SECURITY_ERROR`; see + [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). - **Network and availability failures** (connection refused, TLS errors, a `5xx` from the server, a mid-query drop) are treated as temporary and fed into the reconnect and failover loops. diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 3b3b26ead1..524708276e 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -3,8 +3,9 @@ slug: /connect/clients/nodejs title: Node.js client for QuestDB sidebar_label: Node.js description: - "QuestDB TypeScript and JavaScript Node.js client for high-throughput QWP - ingestion and streaming SQL queries, with pooling, failover, and store-and-forward." + "TypeScript and JavaScript client for QuestDB on Node.js + (@questdb/nodejs-client): QWP ingestion, streaming SQL queries, failover, and + store-and-forward." --- import SfDedupWarning from "../../partials/_sf-dedup-warning.partial.mdx" @@ -73,25 +74,24 @@ complete API from the package root, ships ES module and CommonJS builds, and bundles TypeScript declarations. There are no other supported import paths. The examples on this page are TypeScript ES modules with top-level `await`. -To run them as plain JavaScript, use ES modules (`.mjs` or `"type": "module"`) -and remove type annotations, type-only imports, and TypeScript assertions -such as `as const`. +To run the [quick start](#quick-start) as TypeScript, save its code as +`example.mts`, then run it from the project directory: -## Quick start +```shell +npm install --save-dev tsx +npx tsx example.mts +``` -Connect with one connect string, write two rows, and try to query the ETH-USD -row. Ingestion is asynchronous, so an immediate read may not see it yet. +The `.mts` extension enables ES modules and top-level `await` without changing +`package.json`. Run other examples the same way after supplying any required +configuration. To run them as plain JavaScript instead, use ES modules (`.mjs` +or `"type": "module"`) and remove type annotations, type-only imports, and +TypeScript assertions such as `as const`. -:::note Existing `trades` tables +## Quick start -This example assumes `trades` does not exist yet. If it already exists, QWP -uses its existing designated timestamp column. For the `trades(ts, ...)` schema -in the [PGWire guide](/docs/connect/compatibility/pgwire/nodejs/), -replace the `timestamp` column with `ts` in every `trades` query on this page. -The sender's `at()` calls need no change: they write to the existing -designated timestamp. - -::: +Connect with one connect string, create a table, write two rows, wait until +QuestDB acknowledges them, and query them back. ```typescript import { @@ -101,34 +101,42 @@ import { const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); try { - // Ingest: borrow a sender, add rows, and close() it to flush the rows and - // return the sender to the pool. The underlying connection stays open. - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.00044) - .at(Date.now(), "ms"); - await sender - .table("trades") - .symbol("symbol", "BTC-USD") - .symbol("side", "sell") - .doubleColumn("price", 39269.98) - .doubleColumn("amount", 0.001) - .at(Date.now(), "ms"); - } finally { - await sender.close(); - } - - // Query: borrow a query lease and iterate the result batches. - // QuestDB applies ingested rows asynchronously, so on a first run this - // query can fail with "table does not exist" or return no rows yet. - // See "Read-after-write" below for the polling pattern. const lease = await db.borrowQuery(); try { + // Create the table first, so the query below cannot hit a missing table. + const ddl = await lease.query( + "CREATE TABLE IF NOT EXISTS trades (" + + "symbol SYMBOL, side SYMBOL, price DOUBLE, amount DOUBLE, " + + "timestamp TIMESTAMP) TIMESTAMP(timestamp) PARTITION BY DAY", + ); + await ddl.completion; + + // Ingest: borrow a sender, add rows, publish them, and wait for the ACK. + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.00044) + .at(Date.now(), "ms"); + await sender + .table("trades") + .symbol("symbol", "BTC-USD") + .symbol("side", "sell") + .doubleColumn("price", 39269.98) + .doubleColumn("amount", 0.001) + .at(Date.now(), "ms"); + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence); + } finally { + // Returns the sender to the pool. The connection stays open. + await sender.close(); + } + + // Query. QuestDB applies acknowledged rows asynchronously, so a query + // right after ingestion can still return no rows; see Read-after-write. const query = await lease.query( "SELECT timestamp, symbol, price, amount FROM trades " + "WHERE symbol = 'ETH-USD' LIMIT 10", @@ -156,23 +164,28 @@ What happens: 1. `connectQwpNodeClient()` validates every key of the connect string, then opens one ingestion and one query connection. It rejects if QuestDB is unreachable. -2. `db.borrowSender()` leases a sender. Rows are staged locally until an - auto-flush threshold is reached or the sender is flushed. `close()` on a - borrowed sender flushes its rows and returns it to the pool. -3. `db.borrowQuery()` leases a query connection. `lease.query()` returns a - query handle that is an async iterable of result batches. `batch.rows()` - yields one array per row. `query.completion` resolves when the server - finishes the query. -4. `db.close()` closes the pools, and resolves even if QuestDB has not - acknowledged every row; see +2. `db.borrowQuery()` leases a query connection. `lease.query()` runs one SQL + statement and returns a query handle; its `completion` promise settles when + the statement ends. +3. `db.borrowSender()` leases a sender. Rows are staged locally until an + auto-flush threshold is reached or the sender is flushed. + `waitForAcknowledged()` waits until QuestDB has committed them, and + `close()` returns the sender to the pool. +4. The query handle is an async iterable of result batches, and `batch.rows()` + yields one array per row. +5. `db.close()` closes both pools; see [Closing the pooled client](#closing-the-pooled-client). -If `trades` did not exist, ingestion creates it automatically with a designated -timestamp column named `timestamp`. Timestamps come back as `bigint` -microseconds since the Unix epoch; see -[Reading result values](#reading-result-values) for every type. +Without the `CREATE TABLE`, the first write creates `trades` automatically, +with a designated timestamp column named `timestamp`. A `trades` table that +already exists is left unchanged, and the sender writes to its designated +timestamp. If yours uses the `trades(ts, ...)` schema from the +[PGWire guide](/docs/connect/compatibility/pgwire/nodejs/), replace +`timestamp` with `ts` in the SQL on this page. -To read your own writes reliably, see [Read-after-write](#read-after-write). +Timestamps come back as `bigint` microseconds since the Unix epoch; see +[Reading result values](#reading-result-values) for every type. To wait until a +write is visible to queries, see [Read-after-write](#read-after-write). ## Connecting @@ -266,12 +279,12 @@ const sender = await connectQwpNodeSender( ); try { await sender - .table("trades") + .table("orders") .symbol("symbol", "ETH-USD") .symbol("side", "buy") + .uuidColumn("order_id", "9f1c96b2-54b8-4d85-bb24-e82c6f1ac120") .doubleColumn("price", 2615.54) .doubleColumn("amount", 0.5) - .uuidColumn("order_id", "9f1c96b2-54b8-4d85-bb24-e82c6f1ac120") .at(Date.now(), "ms"); await sender.flush(); } finally { @@ -321,7 +334,18 @@ A QWP connect string has the form `schema::key=value;key=value;`: transports. - **Values** end at `;`. Double a semicolon to include it in a value: `password=p;;ssw;;rd` sets the password to `p;ssw;rd`. The trailing `;` is - optional. + optional. Every key needs a value: `client_id=;` fails with + `value is not set for 'client_id'`. +- **Each key appears once**, except `addr`. Repeating a key fails with + `Duplicate QWP cluster configuration key: ''`, and so does setting a key + and its alias, such as `user` and `username`. +- **No spaces** around the commas in `addr`: `addr=a:9000, b:9000` fails with + `Invalid QWP cluster address entry: ' b:9000'`. + +To add settings to a connect string that comes from configuration, such as +`QDB_CLIENT_CONF`, append only keys that the string does not set already, or +pass the setting as a [typed option](#programmatic-options), which takes +precedence without a duplicate-key error. The Node.js client's parser differs from some other clients in two places: @@ -338,8 +362,12 @@ For every key and its default, see the ### Programmatic options Callbacks, custom agents, and other settings a string cannot express go in the -second argument. When the connect string and typed options set the same option, -the typed value wins: +second argument, a `QwpNodeClientConfigOptions` object. When the connect string +and typed options set the same option, the typed value wins. Credentials and +TLS are the exception: a typed `webSocket.authorization` header cannot be +combined with `token`, `username`, or `password` in the string, and a typed +`webSocket.agent` cannot be combined with `tls_verify` or `tls_roots`. Both +combinations are rejected. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -366,9 +394,26 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { await db.close(); ``` -The other sections are `webSocket` (connection settings shared by both -directions, such as `agent` or `connectTimeoutMs`) and `storeAndForward` -(journal settings, see [Store-and-forward](#store-and-forward)). +The typed `egress` section takes `target`, `zone`, `compression`, +`compressionLevel`, and `maxBatchRows`. The other sections are `webSocket` +(connection settings shared by both directions, such as `agent` or +`connectTimeoutMs`) and `storeAndForward`, the journal settings described +under [Store-and-forward](#store-and-forward). Its fields match the connect +string keys: + +| Typed field | Connect string key | +|---|---| +| `directory` | `sf_dir` | +| `maxBytes` | `sf_max_total_bytes` | +| `maxSegmentBytes` | `sf_max_segment_bytes` | +| `durability` | `sf_durability` | +| `checkpointIntervalMs` | `sf_sync_interval_millis` | +| `appendDeadlineMs` | `sf_append_deadline_millis` | +| `drainOrphans` | `drain_orphans` | +| `maxBackgroundDrainers` | `max_background_drainers` | + +The second argument has no field for `sender_id`; set it in the connect +string. `Sender.fromConfig()` takes `{ log, agent, qwp }` as its second argument, where `qwp` has the sections `webSocket`, `session` (the equivalent of @@ -414,9 +459,10 @@ username cannot contain `:`. ### TLS -The `wss` schema enables TLS and verifies the server certificate against the -CA certificates bundled with Node.js, not the operating system's trust store. A -private CA installed only in the operating system is not trusted. To trust it, +The `wss` schema enables TLS and, by default, verifies the server certificate +against the CA certificates bundled with Node.js, not the operating system's +trust store. A private CA installed only in the operating system is not +trusted. To trust it, set `tls_roots`, or add it for the whole process with the `NODE_EXTRA_CA_CERTS` environment variable, which Node.js reads at startup. Two keys adjust verification, and both are rejected on a plain `ws` string: @@ -433,28 +479,16 @@ To route the connection through an HTTP or SOCKS proxy, pass an agent such as `Sender`). A custom agent owns certificate verification, so it cannot be combined with `tls_verify` or `tls_roots`. -Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS -and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and -authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits -`connect_timeout` when only the latter is set. A timeout in either phase -produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or -`authentication`. - -After the upgrade, a query connection has a separate 5-second deadline for -the initial QWP `SERVER_INFO` frame. Configure it with the typed option -`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not -change it. Expiry produces an ordinary `Error` with the message -`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. - The pooled client reports connection setup failures as the `cause` of a -`QwpPoolResourceError`; see [Connection-level errors](#connection-level-errors). +`QwpPoolResourceError`. For the setup deadlines and the errors they produce, +see [Connection timeouts](#connection-timeouts). ### Unsupported authentication paths | Path | Status | Workaround | |---|---|---| | OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider and has no callback to refresh a token. | Obtain an access token from your identity provider, pass it as `token=...`, and create a new client before the token expires. See [OpenID Connect](/docs/security/oidc/). | -| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and memory-mode senders then fail, while senders with `sf_dir` or background replay keep retrying and buffering (see [Connection-level errors](#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | +| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and senders in default memory mode then fail, while senders with `sf_dir` or in background memory mode keep retrying and buffering (see [Connection-level errors](#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | | Mutual TLS (client certificates) | Not supported. QuestDB does not negotiate client certificates. | Use token or basic authentication over `wss`. | | ILP JWK authentication | Not available for QWP. `auth`, `jwk`, `token_x`, and `token_y` are rejected on `ws`/`wss`. | Use token or basic authentication. | @@ -503,7 +537,7 @@ try { .at(Date.now(), "ms"); } } finally { - // With default options, flushes and returns without waiting for ACKs. + // Flushes and returns the sender to the pool. Does not wait for ACKs. await sender.close(); } } finally { @@ -516,9 +550,14 @@ A long-running producer can keep its borrow for its whole lifetime and call that hold a sender at the same time. `close()` on a borrowed sender flushes its completed rows and returns it to -the pool without waiting for acknowledgements. See -[Closing a sender](#closing-a-sender) for how to wait for them, and for what -happens when a close fails. +the pool. It does not wait for acknowledgements, but in the default memory +mode its flush waits for the reconnect during an outage. See +[Closing a borrowed sender](#closing-a-borrowed-sender) for how long that can +take, how to wait for acknowledgements, and what happens when a close fails. + +After `close()`, every method call or property read on that sender object +throws `QwpClientClosedError`. Don't keep references to a returned sender, for +example in callbacks that can run later. ### Borrowing a query lease @@ -526,7 +565,10 @@ A query lease runs one query at a time. For concurrent queries, borrow one lease per query, up to `query_pool_max`: ```typescript -import { connectQwpNodeClient, type QwpQueryLease } from "@questdb/nodejs-client"; +import { + connectQwpNodeClient, + type QwpQueryLease, +} from "@questdb/nodejs-client"; const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); @@ -558,7 +600,7 @@ try { } ``` -Starting a second query on a lease while one is still active throws +Starting a second query on a lease while one is still active rejects with `a QWP query is already active on this connection`. Always close a lease in `finally`: an unreturned lease holds its connection until `db.close()`. @@ -591,6 +633,47 @@ The other two settings use different locations: When creating a new pooled connection fails, the borrow rejects with `QwpPoolResourceError`, whose `cause` holds the connection error. +`borrowSender()` and `borrowQuery()` take no timeout argument. When the pool +is at its maximum, a borrow waits up to `acquire_timeout_ms` for a connection +to be returned. Opening a new connection is bounded by the +[connection timeouts](#connection-timeouts) of each endpoint, and by the +[failover budget](#query-failover) when query retries are on. To enforce a +shorter deadline, such as a request deadline, race the borrow against a timer +and return a lease that arrives late: + +```typescript +import { connectQwpNodeClient, type QwpClient } from "@questdb/nodejs-client"; + +function borrowQueryWithin(db: QwpClient, timeoutMs: number) { + const borrow = db.borrowQuery(); + let timer: ReturnType | undefined; + const deadline = new Promise((_, reject) => { + timer = setTimeout(() => reject(new Error("borrow timed out")), timeoutMs); + }); + return Promise.race([borrow, deadline]) + .catch((error: unknown) => { + // Return a lease that arrives after the deadline. + borrow.then((lease) => lease.close(), () => undefined); + throw error; + }) + .finally(() => clearTimeout(timer)); +} + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await borrowQueryWithin(db, 2_000); + try { + const query = await lease.query("SELECT count() FROM trades"); + for await (const batch of query) console.log(batch.get(0, 0)); + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + ### Starting while QuestDB is down `connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set @@ -602,7 +685,9 @@ until the first query. import { connectQwpNodeClient } from "@questdb/nodejs-client"; // Resolves immediately, even if QuestDB is not running yet. -const db = await connectQwpNodeClient("ws::addr=localhost:9000;lazy_connect=on;"); +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;lazy_connect=on;", +); try { const sender = await db.borrowSender(); try { @@ -634,18 +719,10 @@ without `lazy_connect` is not enough: the query pool still connects at startup, so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. -Two caveats apply to a lazy start: - -- **A locked journal still stops it.** With `sf_dir`, the client opens its - journal at startup. If another process holds the journal, or a crashed - process left a stale lock, `connectQwpNodeClient()` rejects with - `QwpPoolResourceError` whose `cause` is `QwpReplayStoreLockedError`, with or - without `lazy_connect=on`. See Lock recovery under - [Store-and-forward](#store-and-forward). -- **Batches are not checked against the server's size limit.** Until a sender - has connected once, it cannot know the limit. If a batch can exceed about - 1 MiB, also set `sf_max_segment_bytes`; see Oversized rows under - [Flushing](#flushing). +A lazy start does not cover two cases. A locked store-and-forward journal +still fails startup; see [Lock recovery](#sf-lock-recovery). And until a +sender has connected once, it cannot check batches against the server's size +limit; see [Batch size limits](#batch-size-limits). ### Closing the pooled client @@ -655,9 +732,11 @@ Two caveats apply to a lazy start: ones. - Closes idle senders. Each publishes its remaining rows and waits up to `close_flush_timeout_millis` (5 seconds) for QuestDB to acknowledge them. -- Waits up to 5 seconds, or `acquire_timeout_ms` if lower, for borrowed - senders to be returned. A sender still borrowed after that stays open, and - its owner must `close()` it. +- Waits for borrowed senders to be returned, until 5 seconds after + `db.close()` was called, or `acquire_timeout_ms` if that is lower. Closing + the idle senders counts toward the same deadline. A sender still borrowed + after that stays open: its owner must `close()` it, and the process stays + alive until then. `db.close()` resolves even when an acknowledgement does not arrive in time. Without `sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they @@ -670,6 +749,26 @@ acknowledgement before returning each sender (see [Awaiting acknowledgements](#awaiting-acknowledgements)), or use [store-and-forward](#store-and-forward). +## Concurrency + +Node.js runs your code on one thread, but async functions interleave at every +`await`: + +- **`QwpClient`** is safe to share across your whole application. +- **Senders** are not safe for concurrent producers. A row is built across + several calls, so an `await` between `table()` and `at()` lets another task + add columns to the same row. Give each producer its own sender, borrowed from + the pool, and size `sender_pool_max` to match. +- **Query leases** run one query at a time. Borrow one lease per concurrent + query; `query_pool_max` caps concurrent queries. +- **Worker threads** cannot share clients. Create one client per worker, and + give each worker its own `sender_id` when using store-and-forward. + +Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy +work out of them. Row encoding runs on the event loop too, so one process +ingests at most as fast as one CPU core allows. To go faster, split the stream +across worker threads or processes, each with its own client. + ## Data ingestion @@ -737,23 +836,9 @@ method called before that throws `table name must be set before adding columns`. drops every row staged since the last flush. An awaited `at()` or `atNow()` can also reject because an auto-flush failed -after the row was completed. What happens to the completed rows depends on the -error: - -- `QwpReplayRejectedError`, `QwpIngressNackError`, `QwpReconnectExhaustedError`, - or `QwpReplayStoreError`: the sender has failed permanently, and every row - still staged on it is lost when it closes. Close the sender, borrow or create - a new one, and write those rows again. With `sf_dir`, fix the cause of a - terminal rejection first, because the rejected batch blocks the journal (see - [Ingestion errors](#ingestion-errors)). A full SYMBOL dictionary also fails - the sender permanently; see [Column methods](#column-methods). -- `QwpBatchTooLargeError`, `QwpMemoryReplayAppendTimeoutError`, or - `QwpReplayStoreAppendTimeoutError`: the completed rows stay staged for a - later `flush()` or `close()`, and the sender stays usable. Do not resubmit - them. - -These classes have no common base class, so test for them by name. See -[Flushing](#flushing) and [Ingestion errors](#ingestion-errors). +after the row was completed. Whether the completed rows are still staged, and +what to do next, depends on the error class; see the +[Error handling](#error-handling) table. ### Column methods @@ -772,7 +857,7 @@ does not exist yet: | `longColumn(name, value)`, `intColumn(name, value)` | LONG | Safe-integer `number` or `bigint`. `-9223372036854775808n` stores NULL | | `float32Column(name, value)` | FLOAT | `number` | | `doubleColumn(name, value)`, `floatColumn(name, value)` | DOUBLE | `number` | -| `timestampColumn(name, value, unit)` | TIMESTAMP, or TIMESTAMP_NS with unit `"ns"` | Integer `number` or `bigint`. Unit `"us"` (default), `"ms"`, or `"ns"`; `"ns"` requires a `bigint` | +| `timestampColumn(name, value, unit?)` | TIMESTAMP, or TIMESTAMP_NS with unit `"ns"` | Integer `number` or `bigint`. Unit `"us"` (default), `"ms"`, or `"ns"`; `"ns"` requires a `bigint` | | `dateColumn(name, value)` | DATE | Epoch milliseconds as `number` or `bigint` | | `charColumn(name, value)` | CHAR | One-character `string` (a single UTF-16 code unit) | | `binaryColumn(name, value)` | BINARY | `Uint8Array`, copied when staged | @@ -800,15 +885,52 @@ Names that differ from what you might expect: - `geohashColumn()` takes raw bits only. Base-32 geohash text is accepted by a compiled writer's `geohash()` field. +One row can mix any of these methods. This example creates an `orders` table +with a UUID, INT, LONG, BOOLEAN, and a second TIMESTAMP column: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + const submittedMs = Date.now() - 250; + await sender + .table("orders") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .uuidColumn("order_id", "0b2b6c4e-7c39-4f0e-9d5a-2f8e61c3a7d4") + .doubleColumn("price", 2615.54) // DOUBLE + .doubleColumn("amount", 0.5) + .int32Column("venue_id", 7) // INT + .longColumn("lots", 125n) // LONG + .booleanColumn("is_maker", true) // BOOLEAN + .timestampColumn("submitted_at", submittedMs, "ms") // TIMESTAMP + .at(Date.now(), "ms"); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + :::caution SYMBOL is for bounded sets of values Use SYMBOL for values from a bounded set, such as tickers, sides, or venues. Each sender keeps every distinct SYMBOL value it has sent, across all tables -and columns, in a dictionary that holds at most 2,000,000 values and is not -cleared while the sender lives. Once it is full, every later `at()`, `flush()`, -and `close()` on that sender fails with an `Error`, and the sender's staged -rows are lost. Store unique or high-cardinality values, such as trade or order -IDs, as VARCHAR with `stringColumn()`, or as UUID with `uuidColumn()`. See +and columns, in a dictionary that holds at most 2,000,000 values. The +dictionary is not cleared while the sender lives, and no metric reports its +size. Staging a row never fails because of it: the flush that sends a batch +with a value beyond the limit fails with an `Error`, whether it is an explicit +`flush()`, the auto-flush of an `at()`, or `close()`. The rows stay staged, so +every later flush fails the same way. Call `reset()` to drop them; rows with +values the sender already knows can still be sent. Only a new sender starts +with an empty dictionary: close a standalone sender and create another. A +pooled sender whose `close()` fails with this error is replaced by the pool. +Store unique or high-cardinality values, such as trade or order IDs, as VARCHAR +with `stringColumn()`, or as UUID with `uuidColumn()`. See [Symbol](/docs/concepts/symbol/). ::: @@ -987,9 +1109,9 @@ Every sub-array at the same depth must have the same length, and arrays may have 1 to 32 dimensions. Only DOUBLE arrays can be ingested: `longArrayColumn()` exists for protocol parity, but current servers reject it with `long arrays are not supported, only double arrays`. The rejection is -terminal: it stops the sender, and with `sf_dir` the rejected batch blocks the -journal for every table (see [Store-and-forward](#store-and-forward)). Query -results return arrays as `{ dimensions, values }`; see +terminal; see +[Recovering from a terminal rejection](#recovering-from-a-terminal-rejection). +Query results return arrays as `{ dimensions, values }`; see [Reading result values](#reading-result-values). @@ -1097,9 +1219,15 @@ try { timestamp: Date.now(), }); - // Arrays, iterables, and async iterables. Absent nullable fields store NULL. + // Arrays, iterables, and async iterables. + // Absent nullable fields store NULL. await trades.rows([ - { symbol: "BTC-USD", side: "buy", price: 39269.98, timestamp: Date.now() }, + { + symbol: "BTC-USD", + side: "buy", + price: 39269.98, + timestamp: Date.now(), + }, ]); } catch (error) { if (!(error instanceof QwpWriterRowError)) throw error; @@ -1159,20 +1287,23 @@ default and flushes after the row that crosses the first threshold: | Trigger | Default | Connect-string key | Typed option | |---|---|---|---| | Row count | 1,000 rows | `auto_flush_rows` | `autoFlushRows` | -| Time since the last flush | 100 ms | `auto_flush_interval` | `autoFlushIntervalMs` | +| Time since the last flush, or since the sender was created | 100 ms | `auto_flush_interval` | `autoFlushIntervalMs` | | Estimated buffered bytes | Disabled | `auto_flush_bytes` | `autoFlushBytes` | The interval is checked when a row is added. There is no background timer, so call `flush()` after a burst of rows, or rows staged before an idle period wait for the next row. `auto_flush=off` disables all triggers. `auto_flush_bytes` is -clamped to 90% of the batch size the server advertises. +clamped to 90% of the effective batch limit: the server's limit, or +`sf_max_segment_bytes` when that is lower (see +[Batch size limits](#batch-size-limits)). -What `flush()` waits for depends on the ingestion mode: +What `flush()` waits for depends on the ingestion mode. This page uses these +three names for the modes: | Mode | Enabled by | `flush()` resolves when | During an outage | |---|---|---|---| -| Memory (default) | Neither of the others | The batch is written to the WebSocket, or queued for replay | `flush()` and auto-flushing `at()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | -| Background memory | `initial_connect_retry=async`, or `lazy_connect=on` on the pooled client | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | +| Default memory mode | Neither of the others | The batch is written to the WebSocket, or queued for replay | `flush()`, auto-flushing `at()`, and a borrowed sender's `close()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | +| Background memory mode | `initial_connect_retry=async` or `lazy_connect=on` | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | | Store-and-forward | `sf_dir` | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | In every mode, `flush()` does not wait for QuestDB to acknowledge the rows, @@ -1180,38 +1311,57 @@ unless you set `awaitServerAck`. Unacknowledged batches are kept and replayed after a reconnect. See [Awaiting acknowledgements](#awaiting-acknowledgements) and [Store-and-forward](#store-and-forward). -**Backpressure.** The in-memory replay queue is capped at 128 MiB. When it is -full, publishing waits up to 30 seconds for acknowledgements to free space, -then rejects with `QwpMemoryReplayAppendTimeoutError`. Tune the cap with -`sf_max_total_bytes` and the wait with `sf_append_deadline_millis`; without -`sf_dir` they size the memory queue. Watch `sender.metrics.ingress` -(`memoryReplayUsedBytes`, `totalMemoryReplayBackpressureStalls`) to detect -backpressure before it blocks. `metrics` is available on pooled senders and on -senders from `connectQwpNodeSender()`, not on the standalone `Sender` class. - -**Oversized rows.** When the sender connects, QuestDB advertises the largest -batch it accepts: about 2 MiB on a default server, set by +#### Backpressure + +The in-memory replay queue is capped at 128 MiB. When it is full, publishing +waits up to 30 seconds for acknowledgements to free space, then rejects with +`QwpMemoryReplayAppendTimeoutError`. Tune the cap with `sf_max_total_bytes` and +the wait with `sf_append_deadline_millis`; without `sf_dir` they size the +memory queue. A store-and-forward journal applies the same backpressure and +rejects with `QwpReplayStoreAppendTimeoutError`; see +[Journal capacity](#sf-capacity). + +After an append timeout the batch stays staged and the sender stays usable. +Keep the sender, slow the producer, and call `flush()` again later. Don't write +the rows again, and don't `close()` a borrowed sender while backpressure +persists: its `close()` flushes too, so it can time out the same way, and the +pool discards a borrowed sender whose `close()` fails. In the memory modes, +every batch it held that QuestDB has not acknowledged by the time it closes is +lost with it. + +`sender.metrics.ingress` shows the backlog in every mode: `pendingReplayFrames` +and `pendingReplayBytes` count the published batches that QuestDB has not +acknowledged yet. In the memory modes, `memoryReplayUsedBytes` and +`memoryReplayMaxBytes` show how full the replay queue is, and +`totalMemoryReplayBackpressureStalls` counts publishes that had to wait. +`metrics` is available on pooled senders and on senders from +`connectQwpNodeSender()`, not on the standalone `Sender` class. + +#### Batch size limits + +When the sender connects, QuestDB advertises the largest batch it accepts: +about 2 MiB (2,097,138 bytes) on a default server, set by `http.recv.buffer.size`. A batch must also fit in `sf_max_segment_bytes`, which defaults to 4 MiB with `sf_dir` and applies without `sf_dir` only when -you set it. A row too large to fit in one batch fails the -`flush()`, or the `at()` whose auto-flush sends it, with -`QwpBatchTooLargeError` before anything is sent. The staged rows are kept, so -every later flush fails the same way, and `close()` discards them and rejects -with the same error. Call `reset()` to drop every row staged since the last -flush, then write the other rows again. - -**Batches staged before the first connection.** Until a sender has connected -once, it does not know the server's limit. This applies with -`lazy_connect=on`, with `initial_connect_retry=async`, and to a -store-and-forward restart while QuestDB is down. Batches are then capped only -by `sf_max_segment_bytes`: 4 MiB with `sf_dir`, and no cap without it. A batch +you set it. + +A row too large to fit in one batch fails the `flush()`, or the `at()` whose +auto-flush sends it, with `QwpBatchTooLargeError` before anything is sent. +That batch can never be sent: the staged rows are kept, every later flush fails +the same way, and `close()` discards them and rejects with the same error. +Call `reset()` to drop every row staged since the last flush, then write the +rows again without the oversized one. + +Until a sender has connected once, it does not know the server's limit. This +applies in background memory mode, and to a store-and-forward sender that +restarts while QuestDB is down. Batches are then capped only by +`sf_max_segment_bytes`: 4 MiB with `sf_dir`, and no cap without it. A batch larger than the server's limit passes `flush()` but can never be delivered: the sender keeps reconnecting, and `waitForAcknowledged()` times out. With `sf_dir`, the batch also blocks the journal, so later rows are not delivered and a restarted client fails with `QwpPoolResourceError`. If the client can -start while QuestDB is down and a batch can exceed about 1 MiB, set -`sf_max_segment_bytes=1m`: `2m` is slightly above a default server's limit of -2,097,138 bytes. +start while QuestDB is down, set `sf_max_segment_bytes` below the server's +limit, for example `1m`: `2m` is slightly above a default server's limit. ### Closing a sender @@ -1219,7 +1369,7 @@ start while QuestDB is down and a batch can exceed about 1 MiB, set with a warning. The rest depends on how you created the sender. With transactions on, see also [Transactions](#transactions). -#### Standalone sender +#### Closing a standalone sender `close()` waits up to `close_flush_timeout_millis` (5 seconds by default) for QuestDB to acknowledge every published row, then closes the connection. `0` or @@ -1257,25 +1407,36 @@ try { } ``` -#### Borrowed sender +#### Closing a borrowed sender -`close()` returns the sender to the pool without closing its connection and, -by default, without waiting for acknowledgements. With `awaitServerAck: true` -or `awaitDurableAck: true`, the flush performed by `close()` waits for its -acknowledgement too. To confirm delivery of every row the sender published -before returning it, call `flush()` and then -`waitForAcknowledged(sender.publishedSequence)`; see +`close()` flushes the sender's completed rows and returns it to the pool +without closing its connection. By default it does not wait for +acknowledgements. With `awaitServerAck: true` or `awaitDurableAck: true`, the +flush performed by `close()` waits for its acknowledgement too. To confirm +delivery of every row the sender published before returning it, call +`flush()` and then `waitForAcknowledged(sender.publishedSequence)`; see [Awaiting acknowledgements](#awaiting-acknowledgements). +In the default memory mode, the flush in `close()` behaves like `flush()` +during an outage: it waits for the reconnect, up to +`reconnect_max_duration_millis` (5 minutes by default), then rejects with +`QwpReconnectExhaustedError`, and the rows are lost. A standalone sender's +`close()` is bounded by `close_flush_timeout_millis` instead. `db.close()` does +not wait for such a `close()` to finish, and the process stays alive until it +does. To keep shutdown within a deadline, such as a container's termination +grace period, lower `reconnect_max_duration_millis` to fit it, or use +store-and-forward: with `sf_dir`, `flush()` appends to the journal and +returns, and the rows survive the restart. + When a borrowed sender's `close()` fails, the pool discards the sender and -opens a new one for the next borrow. Because QuestDB reports rejected batches +opens a new one for the next borrow. In the memory modes, the batches that +QuestDB has not acknowledged by the time the discarded sender closes are lost +with it; with `sf_dir`, they stay in its journal. Because QuestDB reports rejected batches asynchronously, a sender can fail after its `close()` already succeeded: the error then surfaces on the next borrower's auto-flushing `at()`, `flush()`, or `close()`, and the pool replaces the sender after that. The next borrower's own staged rows are lost with the failed sender, even rows for other tables: write -them again on a new borrow, as described in -[General usage pattern](#general-usage-pattern). See -[Ingestion errors](#ingestion-errors). +them again on a new borrow. See [Ingestion errors](#ingestion-errors). ### Awaiting acknowledgements @@ -1387,7 +1548,8 @@ try { const pending: { sequence: bigint; offset: bigint }[] = []; const commitAcknowledged = async () => { let offset: bigint | undefined; - while (pending.length > 0 && pending[0].sequence <= sender.acknowledgedSequence) { + const acknowledged = sender.acknowledgedSequence; + while (pending.length > 0 && pending[0].sequence <= acknowledged) { offset = pending.shift()!.offset; } if (offset !== undefined) await commitOffset(offset); @@ -1425,8 +1587,11 @@ try { ``` To commit as acknowledgements arrive instead of after each flush, register -`ingressSession.onProgress`: it receives events whose `kind` is `published`, -`acknowledged`, or `durable-acknowledged`, with the `sequence` they cover. +`ingressSession.onProgress`. It receives `QwpIngressProgressEvent` objects whose +`kind` is `published`, `acknowledged`, or `durable-acknowledged`, with the +`sequence` they cover. Read the watermark from the event, not from the +sender: events can arrive after a borrowed sender was returned to the pool, +and a returned sender throws `QwpClientClosedError` on every access. ### Transactions @@ -1475,7 +1640,8 @@ try { have `commit()`, an alias of `flush()`. The typed option is `transactional: true`. - Closing a standalone sender without calling `flush()` rolls the open - transaction back, with a warning. + transaction back, with a warning. Tables and columns that the rolled-back + batches created remain. - Returning a borrowed sender with `close()` commits instead, because `close()` flushes before returning the sender to the pool. `reset()` does not prevent this: it drops only rows staged since the last flush, not the batches already @@ -1492,18 +1658,27 @@ exits. Setting `sf_dir` turns on a disk journal instead: every batch is appended to the journal before it is sent, a background drainer sends it in order, and acknowledged segments are deleted. -Before ingesting, create a deduplicated table while QuestDB is reachable, for -example with `lease.query()` as in -[DDL and DML statements](#ddl-and-dml-statements). If the table does not exist -when the first batch arrives, QWP creates it without deduplication, and -replayed batches can then insert duplicate rows. Use both the event timestamp -and a stable, source-assigned trade ID as upsert keys: distinct trades can -share a millisecond timestamp, symbol, and side. +A frame appended to the journal but not acknowledged before a crash is sent +again, so delivery is at least once: + + + +If the table does not exist when the first batch arrives, QWP creates it +without deduplication, and replayed batches can then insert duplicate rows. +Create the table as part of a deployment or schema migration, before any +sender writes to it. A sender that starts while QuestDB is down, as in the +example below, delivers its journal as soon as QuestDB is reachable, which can +be before your own startup code gets to run DDL. If the table may already +exist without deduplication, enable it with +`ALTER TABLE ... DEDUP ENABLE UPSERT KEYS(...)`, which is safe to run again. + +Use both the event timestamp and a stable, source-assigned trade ID as upsert +keys: distinct trades can share a millisecond timestamp, symbol, and side. Store the trade ID as VARCHAR, not SYMBOL: every trade has its own ID, and SYMBOL is for [bounded sets of values](#column-methods). ```questdb-sql -CREATE TABLE trades_sf ( +CREATE TABLE IF NOT EXISTS trades_sf ( timestamp TIMESTAMP, trade_id VARCHAR, symbol SYMBOL, @@ -1519,13 +1694,26 @@ event. The following values represent one source event; do not regenerate them when retrying it: ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { + connectQwpNodeClient, + QwpPoolResourceError, + QwpReplayStoreLockedError, +} from "@questdb/nodejs-client"; const db = await connectQwpNodeClient( "ws::addr=localhost:9000;" + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + "sf_durability=append;lazy_connect=on;", -); +).catch((error: unknown) => { + // Another process holds the journal, or a crash left a stale lock. + if ( + error instanceof QwpPoolResourceError && + error.cause instanceof QwpReplayStoreLockedError + ) { + console.error("the journal is locked; see Lock recovery"); + } + throw error; +}); const event = { tradeId: "trade-12345", timestampMs: 1723000000000 }; try { const sender = await db.borrowSender(); @@ -1549,10 +1737,10 @@ try { ``` With a journal, the sender keeps accepting rows while QuestDB is unreachable, -subject to journal capacity and backpressure as described below. It retries the -connection indefinitely once it has connected, and a new sender opened on the -same directory replays what the previous process left behind, once it can take -over the directory's lock (see Lock recovery below). +subject to [journal capacity](#sf-capacity). It retries the connection +indefinitely once it has connected, and a new sender opened on the same +directory replays what the previous process left behind, once it can take over +the directory's lock (see [Lock recovery](#sf-lock-recovery)). - **Layout.** A standalone `Sender` journals into `/`. A pooled client uses one directory per pooled sender: @@ -1563,98 +1751,81 @@ over the directory's lock (see Lock recovery below). drains any of its own `-` journals that no pooled sender holds, such as those left by a larger pool before a restart, without `drain_orphans`. -- **Durability.** `sf_durability` sets how the journal reaches the disk. It is - unrelated to memory mode, which means running without `sf_dir`. `memory` - (the connect-string default) relies on the operating system to write the +- **Durability.** `sf_durability` sets how the journal reaches the disk. Its + `memory` value is unrelated to the memory ingestion modes. `memory` (the + connect-string default) relies on the operating system to write the journal, which survives a process crash but not a power loss. `periodic` checkpoints in the background every `sf_sync_interval_millis` (5 seconds). `append` makes every append durable before `flush()` resolves, which adds a disk sync to every flush: on a producer that flushes often, prefer `periodic` or larger batches. -- **Capacity.** With `sf_dir`, `sf_max_total_bytes` (10 GiB by default) is a - journal size target, not a hard disk limit. Transaction-closing batches can - reserve extra segments so a full journal does not block the commit needed - to release space. Segment reservations can reach roughly twice the target, - depending on segment rounding; retained symbol dictionaries and other - metadata take additional space. Provision headroom for every sender and - monitor actual disk usage. Without `sf_dir`, the key caps the in-memory - replay queue instead. -- **Backpressure.** When an append cannot fit within the journal's capacity - allowances, publishing waits up to `sf_append_deadline_millis` (30 seconds) - for acknowledgements to free space, then rejects with - `QwpReplayStoreAppendTimeoutError`. - **Startup.** To start the pooled client while QuestDB is down, see - [Starting while QuestDB is down](#starting-while-questdb-is-down), including - its caveats about locked journals and batch sizes. A standalone `Sender` - needs only `initial_connect_retry=async`. With the default `off`, the first + [Starting while QuestDB is down](#starting-while-questdb-is-down). A + standalone `Sender` needs only `initial_connect_retry=async` or + `lazy_connect=on`. With the default `initial_connect_retry=off`, the first connection must succeed. -- **Lock recovery.** The Node.js client locks a journal directory with a - `.lock.owner` directory inside it, which records the owner's host name and - process ID, instead of an operating-system file lock. After a crash, a new - sender takes over automatically only when the owner ran on the same host and - its process ID is no longer in use. Otherwise opening the journal fails with - `QwpReplayStoreLockedError`. For the pooled client, `connectQwpNodeClient()` - rejects with it as the `cause` of a `QwpPoolResourceError`, even with - `lazy_connect=on`, so the whole client fails to start, queries included. A - standalone `Sender` rejects on `connect()`. This is common in containers: - the application usually runs as process ID 1, which is in use again after a - restart, and a replacement container usually has a different host name. Once - you have - verified that the previous owner has exited and no process is using the slot, - remove its stale `//.lock.owner` directory and restart the sender. - Here `` is ``, or `-` for a pooled sender. If - startup still reports `QwpReplayStoreLockedError`, also inspect - `/.slot-locks/.lock.owner`: this short-lived guard can survive a - crash during lock acquisition or quarantine. Remove that specific owner - directory only after verifying its owner has exited. Never delete the shared - `.slot-locks` directory or another slot's locks. Automate this cleanup only - where the deployment guarantees that the previous owner has exited before a - new one starts, for example a single replica that uses the `Recreate` update - strategy and a `ReadWriteOnce` volume. A startup step can then remove the - stale owner directories of the client's own slots before it creates the - client. Anywhere two processes can overlap, recover manually. -- **Rejected batches.** A batch that QuestDB rejects terminally, such as one - with a value of the wrong type for an existing column, stays at the head of - the journal. Every new sender on that directory, including the pool's - replacement for a failed pooled sender and the same client after a restart, - sends it again and fails the same way, with `QwpReplayRejectedError`, or - `QwpReplayStoreError` while the failed journal closes. Ingestion through the - client therefore stops for every table, not only the table in the rejected - batch. Treat a terminal rejection as an outage: alert on it from - `onSenderError`, then fix the cause or move the journal aside, as described - in [Ingestion errors](#ingestion-errors). `longArrayColumn()` triggers this - on every current server, which rejects LONG arrays terminally. +- **Rejected batches.** A batch that QuestDB rejects terminally stays at the + head of the journal and stops ingestion through the client, for every table, + until you act; see + [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection). - **Orphans.** With `drain_orphans=on`, a sender also adopts and drains journals with other `sender_id` values left under the same `sf_dir` by processes that crashed, up to `max_background_drainers` (4) at a time. +- **Other clients.** Don't let a client in another language, such as Java, use + an `sf_dir` while a Node.js client runs on it. Those clients lock journals + with operating-system file locks, and neither kind of client sees the + other's locks, so either could open or drain a journal that the other is + writing and corrupt it. The journal format is shared: once every Node.js + client on the directory has stopped, another client can open the journals + they left behind. -A frame appended to the journal but not acknowledged before a crash is sent -again, so delivery is at least once: - - - -The `trades_sf` keys identify a trade without collapsing distinct trades -that share a millisecond timestamp, symbol, and side. For an existing table -that already has a stable `trade_id` column, enable deduplication with -`ALTER TABLE trades_sf DEDUP ENABLE UPSERT KEYS(timestamp, trade_id);`. Deduplication recognizes a replayed row only when it carries the same designated timestamp and trade ID, so reuse event values on application retries instead of calling `atNow()` or generating a new ID. See [Deduplication](/docs/concepts/deduplication/) for choosing keys. -:::warning Share a journal directory only among Node.js clients - -Node.js clients can share an `sf_dir`. Their locks keep each journal to one -process at a time: a second process that opens a journal in use fails with -`QwpReplayStoreLockedError`. Clients in other languages, such as Java, use -operating-system file locks instead, and neither kind of client sees the -other's locks. Such a client must not use the directory while any Node.js -client is running on it: either client could open a journal that the other is -writing, or drain it as an orphan, and corrupt it. Stop every Node.js client on -the directory first. The journal format is shared, so the other client can then -open the journals that the Node.js clients left behind. - -::: +#### Journal capacity {#sf-capacity} + +With `sf_dir`, `sf_max_total_bytes` (10 GiB by default) is a journal size +target, not a hard disk limit. Transaction-closing batches can reserve extra +segments so a full journal does not block the commit needed to release space. +Segment reservations can reach roughly twice the target, depending on segment +rounding; retained symbol dictionaries and other metadata take additional +space. Provision headroom for every sender and monitor actual disk usage. +Without `sf_dir`, the key caps the in-memory replay queue instead. + +When an append cannot fit within these allowances, publishing waits up to +`sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, +then rejects with `QwpReplayStoreAppendTimeoutError`; see +[Backpressure](#backpressure) for what to do next. +`sender.metrics.ingress.pendingReplayBytes` reports how much of the journal +QuestDB has not acknowledged yet. + +#### Lock recovery {#sf-lock-recovery} + +The Node.js client locks a journal directory with a `.lock.owner` directory +inside it, which records the owner's host name and process ID, instead of an +operating-system file lock. After a crash, a new sender takes over +automatically only when the owner ran on the same host and its process ID is +no longer in use. Otherwise opening the journal fails with +`QwpReplayStoreLockedError`: + +- The pooled client opens its senders' journals when it starts, so + `connectQwpNodeClient()` rejects with `QwpPoolResourceError` whose `cause` is + `QwpReplayStoreLockedError`, even with `lazy_connect=on`, and the whole + client fails to start, queries included. With `sender_pool_min=0`, the first + `borrowSender()` rejects instead. +- A standalone `Sender` rejects on `connect()`. + +This is common in containers: the application usually runs as process ID 1, +which is in use again after a restart, and a replacement container usually has +a different host name. Once you have verified that the previous owner has +exited and no process is using the slot, remove its stale +`//.lock.owner` directory and start the client again. `` is +``, or `-` for a pooled sender. If startup still +fails, a guard under `/.slot-locks` may have survived the crash too; +[Node.js lock recovery](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery) +covers it, and when this cleanup can be automated. For all tuning options, see [Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/) @@ -1672,19 +1843,44 @@ configured. By default QuestDB acknowledges a batch when it is committed to the primary's write-ahead log. With `request_durable_ack=on`, the acknowledgement watermark advances only after the batch is uploaded to the replication object store, so -`waitForAcknowledged()` confirms durable upload: +`waitForAcknowledged()` confirms durable upload. To make every `flush()` wait +for durability, also set the typed option `awaitDurableAck`: -```text -wss::addr=db.example.com:9000;token=YOUR_TOKEN;request_durable_ack=on; +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); +const db = await connectQwpNodeClient( + `wss::addr=db.example.com:9000;token=${token};request_durable_ack=on;`, + // Every flush() waits until its batch is in the object store. + { sender: { awaitDurableAck: true } }, +); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} ``` -To make every flush wait for durability, add the typed option -`{ sender: { awaitDurableAck: true } }`. If the server does not support durable -acknowledgement, connecting fails with `QwpDurableAckUnavailableError`, which -the pooled client reports as the `cause` of a `QwpPoolResourceError`. +If the server does not support durable acknowledgement, connecting fails with +`QwpDurableAckUnavailableError`, which the pooled client reports as the +`cause` of a `QwpPoolResourceError`. -Background senders with `initial_connect_retry=async` or `lazy_connect=on` -instead keep retrying and emit `durable-ack-unavailable` +Senders in background memory mode instead keep retrying and emit +`durable-ack-unavailable` [connection events](#connection-events). A store-and-forward sender does the same when reconnecting after its first successful connection. Monitor these events and buffer usage: successful background startup does not confirm that @@ -1757,7 +1953,9 @@ try { } } const completion = await query.completion; - console.log("rows:", completion.kind === "result-end" && completion.totalRows); + if (completion.kind === "result-end") { + console.log("rows:", completion.totalRows); + } } catch (error) { if (!(error instanceof QwpEgressQueryError)) throw error; console.error(`query failed: status=${error.status} ${error.message}`); @@ -1769,9 +1967,8 @@ try { } ``` -`lease.query(sql, options)` sends the query and resolves with a query handle. -Iterating it with `for await` yields `QwpResultBatch` objects, and the handle's -`completion` promise settles when the query ends: +`lease.query(sql, options?)` sends the query and resolves with a +`QwpEgressQuery` handle. The options are a `QwpEgressQueryOptions` object: | Query option | Default | Purpose | |---|---|---| @@ -1781,6 +1978,19 @@ Iterating it with `for await` yields `QwpResultBatch` objects, and the handle's | `autoCredit` | `true` | Replenish the credit window as batches are consumed. | | `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | +There is no per-query failover setting; see [Query failover](#query-failover). +The `QwpEgressQuery` handle has these members: + +| Member | Purpose | +|---|---| +| `for await (const batch of query)` | Yields `QwpResultBatch` objects in order. | +| `completion` | `Promise` that settles when the query ends. | +| `cancel()` | Asks QuestDB to stop the query; see [Cancellation and timeouts](#cancellation-and-timeouts). | +| `awaitCompletion(timeoutMs)` | Resolves `false` if the query is still running after `timeoutMs`. | +| `isDone()` | Whether the query has ended. | +| `grantCredit(bytes)` | Adds flow-control credit when `autoCredit` is `false`. | +| `requestId` | The `bigint` that numbers queries on this connection. | + Iteration and `completion` reject with the same error when the query fails. Consume the result through `for await`, or `await query.completion` directly for statements that return no rows. @@ -1789,7 +1999,11 @@ A `QwpResultBatch` has: - `rowCount` and `columns`: an array of `{ name, type, values, scale?, precisionBits? }`, where `values` holds one entry per row and `type` is the numeric QWP type - code (compare it with the exported `QWP_COLUMN_TYPE` constants). + code. Compare it with the exported `QWP_COLUMN_TYPE` constants: `BOOLEAN`, + `BYTE`, `SHORT`, `CHAR`, `INT`, `LONG`, `FLOAT`, `DOUBLE`, `SYMBOL`, + `VARCHAR`, `TIMESTAMP`, `TIMESTAMP_NANOS`, `DATE`, `UUID`, `LONG256`, + `GEOHASH`, `IPV4`, `BINARY`, `DOUBLE_ARRAY`, `LONG_ARRAY`, `DECIMAL64`, + `DECIMAL128`, and `DECIMAL256`. - `rows()`: a generator that yields one array of values per row. - `get(rowIndex, columnIndex)`: one value. - `batchSequence`: the batch's position in the result, starting at `0n`. @@ -1868,10 +2082,14 @@ const ipv4ToString = (value: number) => // Any row or value to JSON, with bigint values as decimal strings const toJson = (value: unknown) => - JSON.stringify(value, (_key, v) => (typeof v === "bigint" ? v.toString() : v)); + JSON.stringify(value, (_key, v) => + typeof v === "bigint" ? v.toString() : v, + ); console.log(toDate(1723000000000000n).toISOString()); -console.log(uuidToString({ low: 13485158461794337056n, high: 11465204444048149893n })); +console.log( + uuidToString({ low: 13485158461794337056n, high: 11465204444048149893n }), +); console.log(ipv4ToString(-1062731775)); console.log(toJson([1723000000000000n, "ETH-USD", 2615.54])); ``` @@ -1923,17 +2141,17 @@ try { | `setLong(index, value)` | LONG (`number` or `bigint`) | | `setFloat(index, value)` | FLOAT | | `setDouble(index, value)` | DOUBLE | -| `setDate(index, millis)` | DATE | -| `setTimestampMicros(index, micros)` | TIMESTAMP | -| `setTimestampNanos(index, nanos)` | TIMESTAMP_NS | +| `setDate(index, millis)` | DATE (`number` or `bigint`) | +| `setTimestampMicros(index, micros)` | TIMESTAMP (`number` or `bigint`) | +| `setTimestampNanos(index, nanos)` | TIMESTAMP_NS (`number` or `bigint`) | | `setVarchar(index, value)` | VARCHAR, STRING, and SYMBOL comparisons. `null` binds NULL | | `setUuid(index, value)` or `setUuid(index, low, high)` | UUID, as a canonical string or two 64-bit halves. `null` binds NULL | -| `setLong256(index, w0, w1, w2, w3)` | LONG256, least significant word first | -| `setGeohash(index, precisionBits, value)` | GEOHASH | -| `setDecimal64(index, scale, unscaled)` | DECIMAL64 | -| `setDecimal128(index, scale, low, high)` | DECIMAL128 | -| `setDecimal256(index, scale, w0, w1, w2, w3)` | DECIMAL256 | -| `setNull(index, type)` | A typed NULL. `type` is one of the scalar `QwpBindType` values in `QWP_COLUMN_TYPE`, not every column type: BINARY, IPv4, arrays, STRING, and SYMBOL are excluded (bind text as VARCHAR). | +| `setLong256(index, w0, w1, w2, w3)` | LONG256, least significant word first. Each word is a `number` or `bigint` | +| `setGeohash(index, precisionBits, value)` | GEOHASH. `value` is a `number` or `bigint` | +| `setDecimal64(index, scale, unscaled)` | DECIMAL64. `unscaled` is a `number` or `bigint` | +| `setDecimal128(index, scale, low, high)` | DECIMAL128. Each half is a `number` or `bigint` | +| `setDecimal256(index, scale, w0, w1, w2, w3)` | DECIMAL256. Each word is a `number` or `bigint` | +| `setNull(index, type)` | A typed NULL. `type` is a `QWP_COLUMN_TYPE` constant for a scalar type, for example `setNull(1, QWP_COLUMN_TYPE.DOUBLE)`. The TIMESTAMP_NS constant is `TIMESTAMP_NANOS`. BINARY, IPv4, arrays, and SYMBOL are excluded (bind text as VARCHAR). | | `setNullDecimal64/128/256(index, scale)`, `setNullGeohash(index, precisionBits)` | NULL decimals and geohashes, which carry a scale or precision | There is no setter for BINARY, IPv4, or arrays. Bind IPv4 as a string and cast @@ -1978,7 +2196,8 @@ try { try { const statements = [ "CREATE TABLE IF NOT EXISTS fills (" + - "timestamp TIMESTAMP, symbol SYMBOL, side SYMBOL, price DOUBLE, amount DOUBLE" + + "timestamp TIMESTAMP, symbol SYMBOL, side SYMBOL, " + + "price DOUBLE, amount DOUBLE" + ") TIMESTAMP(timestamp) PARTITION BY DAY", "INSERT INTO fills VALUES (now(), 'ETH-USD', 'buy', 2615.54, 0.5)", "UPDATE fills SET amount = 0.6 WHERE symbol = 'ETH-USD'", @@ -2007,10 +2226,11 @@ try { | `"result-end"` | Queries that return rows | `totalRows` (`bigint`) | | `"exec-done"` | DDL and DML | `rowsAffected` (`bigint`, rows written by an `INSERT`), `operationType` (QuestDB's numeric statement type) | -Only `INSERT` reports a row count in `rowsAffected`. For an `UPDATE` on a WAL -table, the default, it currently holds a transaction number rather than the -number of rows changed, so do not use it to check whether an `UPDATE` matched -any rows. DDL reports `0`, except `TRUNCATE`, which currently reports +Only `INSERT` reliably reports a row count in `rowsAffected`. For an `UPDATE` +on a WAL table, the default, it currently holds a transaction number rather +than the number of rows changed, so do not use it to check whether an `UPDATE` +matched any rows. On a non-WAL table, `UPDATE` reports the rows changed. DDL +reports `0`, except `TRUNCATE`, which currently reports `18446744073709551615n`. Statements run in order on one lease, because each is awaited before the next @@ -2074,7 +2294,10 @@ try { await query.completion; if (!visible) { await new Promise((resolve) => - setTimeout(resolve, Math.min(100, Math.max(0, deadline - Date.now()))), + setTimeout( + resolve, + Math.min(100, Math.max(0, deadline - Date.now())), + ), ); } } @@ -2087,8 +2310,9 @@ try { } ``` -SQL errors now surface instead of being retried. Do not replace the poll with a -fixed sleep: the apply latency varies with load. +Because the table already exists, an SQL error fails the loop at once instead +of being retried; the loop retries only while the row is not yet visible. Do +not replace the poll with a fixed sleep: the apply latency varies with load. ### Cancellation and timeouts @@ -2120,9 +2344,12 @@ QuestDB checks for a cancel between result batches, while it is sending. Set a and stops at the next batch after a cancel, a deadline, or an early exit. Start with 1 MiB; the client replenishes it as your loop consumes batches. Without a credit window, QuestDB streams as fast as the network allows: it may -send the whole result before it reads a cancel, so a cancelled query can -complete normally, and returning the lease after an early exit waits while the -rest of the result streams. A credit window does not interrupt expensive work +send the whole result before it reads a cancel, and the client holds whatever +arrived. A cancelled query can then complete normally, and returning the lease +after an early exit waits while the rest of the result streams. If draining +that backlog takes longer than `query_close_timeout_ms`, the cancel fails with +`QwpEgressQueryCancelTimeoutError` and the connection is closed, even while +your loop keeps consuming. A credit window does not interrupt expensive work before the next batch is ready. ```typescript @@ -2145,7 +2372,9 @@ try { } } catch (error) { if (!(error instanceof QwpEgressQueryTimeoutError)) throw error; - console.warn(`query ${error.requestId} timed out after ${error.timeoutMs} ms`); + console.warn( + `query ${error.requestId} timed out after ${error.timeoutMs} ms`, + ); } finally { // Waits for the cancellation to drain before the lease is reused. await lease.close(); @@ -2155,8 +2384,52 @@ try { } ``` +To stop a result after enough rows, call `cancel()` and keep iterating until +the result ends: + +```typescript +import { + connectQwpNodeClient, + QWP_STATUS, + QwpEgressQueryError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM trades", { + initialCredit: 1024 * 1024, + }); + let rows = 0; + let cancelled = false; + try { + for await (const batch of query) { + rows += batch.rowCount; + if (rows >= 10_000 && !cancelled) { + cancelled = true; + // Keep iterating: the result ends a few batches later. + await query.cancel(); + } + } + await query.completion; + } catch (error) { + const wasCancelled = + error instanceof QwpEgressQueryError && + error.status === QWP_STATUS.CANCELLED; + if (!wasCancelled) throw error; + } + console.log(`read ${rows} rows`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + After a query ends early, its connection stays busy until QuestDB confirms the -cancellation, and another `query()` on the same lease throws +cancellation, and another `query()` on the same lease rejects with `a QWP query is already active on this connection`. Return the lease with `close()` and borrow a new one for the next query. `close()` waits up to `query_close_timeout_ms` (5 seconds) for QuestDB to confirm the cancellation. @@ -2197,7 +2470,8 @@ try { ``` With `autoCredit: false`, call `query.grantCredit(bytes)` yourself. To cap the -rows in each batch, set `max_batch_rows` (1 to 1,048,576). A credit window +rows in each batch, set `max_batch_rows` (1 to 1,048,576), or the typed option +`egress: { maxBatchRows }`; it applies to every query. A credit window also lets QuestDB stop at the next batch after a cancel, a deadline, or an early exit; see [Cancellation and timeouts](#cancellation-and-timeouts) for what your loop must do after `cancel()`. @@ -2282,18 +2556,26 @@ have the details and examples: | Error | Surfaces from | State afterwards | What to do | |---|---|---|---| | `TypeError`, `RangeError`, or `Error` from local validation | The column method or `at()` that staged the value | The row in progress is discarded; the sender stays usable | Fix the value and write the row again | -| `QwpBatchTooLargeError` | `flush()`, or the `at()` whose auto-flush sends the batch | The staged rows are kept, and every later flush fails the same way | Call `reset()`, then write the other rows again; see [Flushing](#flushing) | -| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, or an auto-flushing `at()` | The batch stays staged; the sender stays usable | Slow the producer and flush again later; see [Flushing](#flushing) | +| `QwpBatchTooLargeError` | `flush()`, the `at()` whose auto-flush sends the batch, or `close()` | The batch can never be sent: the staged rows are kept, every later flush fails the same way, and `close()` discards them | Call `reset()`, then write the rows again without the oversized one; see [Batch size limits](#batch-size-limits) | +| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, an auto-flushing `at()`, or `close()` | The batch stays staged; the sender stays usable | Keep the sender and flush again later. Don't write the rows again, and don't close the sender while backpressure persists; see [Backpressure](#backpressure) | | Retriable server rejection | `onSenderError` | The client resends the batch; repeated rejections become terminal | Monitor; no action needed per rejection | -| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls, and with `sf_dir` also `QwpReplayStoreError` | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Ingestion errors](#ingestion-errors) | +| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection) | | `QwpReconnectExhaustedError` on a sender | `onError` with `terminal: true`, then the next `flush()`, `at()`, or `close()` | The sender has failed; unsent rows are lost | Borrow a new sender; see [Ingestion reconnect](#ingestion-reconnect) | -| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See Lock recovery under [Store-and-forward](#store-and-forward) | +| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See [Lock recovery](#sf-lock-recovery) | | `QwpPoolResourceError` with another `cause` | `connectQwpNodeClient()`, `borrowSender()`, or `borrowQuery()` | No connection was opened | Unwrap `cause`; see [Connection-level errors](#connection-level-errors) | | `QwpEgressQueryError` | Query iteration and `completion` | The lease stays usable | Fix the SQL or the bind values | | `QwpEgressQueryTimeoutError`, `QwpEgressQueryAbandonedError` | Query iteration and `completion` | The lease is busy until QuestDB confirms the cancellation | Close the lease and borrow a new one | -| `QwpEgressQueryCancelTimeoutError` | `completion` | The connection is closed | Close the lease and borrow a new one | +| `QwpEgressQueryCancelTimeoutError` | Query iteration and `completion` | The connection is closed | Close the lease and borrow a new one | | `QwpReconnectExhaustedError` on a query | Query iteration and `completion` | The lease stays failed, even after QuestDB recovers | Close the lease and borrow a new one; see [Query failover](#query-failover) | +`QwpReplayStoreAppendTimeoutError` and the other store-and-forward journal +errors extend `QwpReplayStoreError`, so test for the specific classes before +the base class, or branch on `error.retryable`. `true` means the failure is +temporary and the sender stays usable. `false` means the journal itself can no +longer be used, for example `QwpReplayStoreCorruptionError` or +`QwpReplayStoreLockLostError`, and the sender has failed. The other classes in +the table have no common base class: test each with `instanceof`. + ### Ingestion errors Ingestion reports errors in two ways: @@ -2317,10 +2599,14 @@ import { const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { ingressSession: { onSenderError: (error: QwpSenderError) => { - const status = error.serverStatusByte?.toString(16); + // serverStatusByte is absent for client-side errors. + const status = + error.serverStatusByte === undefined + ? "none" + : `0x${error.serverStatusByte.toString(16)}`; console.error( `rejected [${error.category}, policy=${error.appliedPolicy}, ` + - `status=0x${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + + `status=${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + error.serverMessage, ); if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { @@ -2390,12 +2676,26 @@ Retriable rejections of symbol-dictionary catch-up frames are also exempt. The six `on_*_error` connect-string keys are accepted but not applied by this client. -**After a terminal server rejection**, the sender is permanently failed. An +Handling notes: + +- **Message stability.** `serverMessage` is free-form English text from the + server. Its wording can change between releases: branch on `category`, not on + the text. +- **Sensitive data.** Server messages can contain column names and values. + Treat them as untrusted input, and redact them before sending them to + third-party error trackers or showing them to end users. +- **Correlation.** There is no server-side request ID. Correlate with the frame + sequence range, `tableName`, and `detectedAtMs`. + +#### Recovering from a terminal rejection + +After a terminal server rejection, the sender is permanently failed. An already-pending `waitForAcknowledged()` for the rejected batch can reject with `QwpIngressNackError`. Once the terminal failure is latched, new calls to `waitForAcknowledged()`, `flush()`, or `close()` reject with `QwpReplayRejectedError`, whose `status` and message repeat the server's. -Error handlers must allow either class depending on timing. +Error handlers must allow either class depending on timing. Writing LONG arrays +with `longArrayColumn()` triggers a terminal rejection on every current server. Close the sender and create a new one. A pooled sender is replaced automatically after the `close()` that reports the error. What happens to the @@ -2408,24 +2708,14 @@ rejected batch depends on the mode: - **With store-and-forward**, the rejected batch stays at the head of the journal. Every new sender on that directory, including the pool's replacement sender and the same client after a restart, sends it again and - fails the same way. Pooled borrows keep getting that journal, so ingestion - through the client stops for every table, not only the table in the - rejected batch, until you act. Fix the cause so that QuestDB accepts the - batch, for example by adjusting the table schema, or stop the process and - move the journal directory aside. Moving it aside discards every + fails the same way, with `QwpReplayRejectedError`. Pooled borrows keep + getting that journal, so ingestion through the client stops for every table, + not only the table in the rejected batch, until you act. Treat it as an + outage and alert on it from `onSenderError`. Fix the cause so that QuestDB + accepts the batch, for example by adjusting the table schema, or stop the + process and move the journal directory aside. Moving it aside discards every unacknowledged batch in it, not only the rejected one. -Handling notes: - -- **Message stability.** `serverMessage` is free-form English text from the - server. Its wording can change between releases: branch on `category`, not on - the text. -- **Sensitive data.** Server messages can contain column names and values. - Treat them as untrusted input, and redact them before sending them to - third-party error trackers or showing them to end users. -- **Correlation.** There is no server-side request ID. Correlate with the frame - sequence range, `tableName`, and `detectedAtMs`. - ### Query errors Query errors reject both the `for await` iteration and `completion`: @@ -2476,7 +2766,10 @@ connection). The lease remains usable after a `QwpEgressQueryError`. The `QWP_STATUS` export names these codes, for example `QWP_STATUS.PARSE_ERROR`, so code can compare against constants instead of -numbers. +numbers. A status alone does not separate a client mistake from a server +fault: `0x06` covers both bind values that cannot be converted and execution +failures, and `0x0b` covers both the server-side query timeout and memory +limits. Other query errors: @@ -2553,24 +2846,42 @@ An authentication rejection (HTTP 401 or 403) is terminal before a sender's first successful connection and for query connections. It stops the endpoint walk because credentials are assumed to be shared across the cluster. -After a successful connection, regular senders with `sf_dir` or background -memory replay (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying +After a successful connection, regular senders with `sf_dir` or in background +memory mode (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying authentication rejections indefinitely. This lets buffered data drain once -server-side authentication is restored. Memory-mode senders without those -settings, and orphan drainers, do not have this exception. See +server-side authentication is restored. Senders in default memory mode, and +orphan drainers, do not have this exception. See [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) for how other clients behave. Endpoints in error messages have any embedded credentials removed. +#### Connection timeouts + +Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS +and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and +authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits +`connect_timeout` when only the latter is set. A timeout in either phase +produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or +`authentication`. + +After the upgrade, a query connection has a separate 5-second deadline for +the initial QWP `SERVER_INFO` frame. Configure it with the typed option +`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not +change it. Expiry produces an ordinary `Error` with the message +`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. + ### Logging The client writes its own messages to the console by default, at the `error`, `warn`, and `info` levels. To route a sender's messages, such as warnings about -rows discarded on close, through your logger, pass a `(level, message)` +rows discarded on close, through your logger, pass a `QwpSenderLogger` function: `{ sender: { log } }` as the second argument of -`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. The function -also receives `debug` messages, one per staged row, so filter by level. +`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. Its signature +is `(level: "error" | "warn" | "info" | "debug", message: string | Error)`, so +convert `message` with `String()` if your logger takes strings only. The +function also receives `debug` messages, one per staged row, so filter by +level. Rejected batches and session errors go to `ingressSession.onSenderError` and `ingressSession.onError`. Their defaults log to the console, so replace both to @@ -2649,19 +2960,17 @@ jitter, then resends every unacknowledged batch: |---|---|---| | `reconnect_initial_backoff_millis` | `100` | First retry delay. | | `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | -| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in memory mode, without `sf_dir` or background replay. `0` removes the limit. | +| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in default memory mode. `0` removes the limit. | | `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | Whether the sender gives up depends on the mode (see [Flushing](#flushing)): -- **Memory mode**, without `sf_dir` or background replay, retries for up to - `reconnect_max_duration_millis` per outage. When the budget runs out, the - sender fails permanently with `QwpReconnectExhaustedError`, and its unsent - rows are lost. The Java reference client retries indefinitely in this mode - instead. -- **Background memory mode** (`initial_connect_retry=async`, or - `lazy_connect=on` on the pooled client) and **store-and-forward** - (`sf_dir`) retry indefinitely. +- **Default memory mode** retries for up to `reconnect_max_duration_millis` + per outage. When the budget runs out, the sender fails permanently with + `QwpReconnectExhaustedError`, and its unsent rows are lost. The Java + reference client retries indefinitely in this mode instead. +- **Background memory mode** (`initial_connect_retry=async` or + `lazy_connect=on`) and **store-and-forward** (`sf_dir`) retry indefinitely. Setting any `reconnect_*` key also makes a sender's first connection retry within the budget, as if `initial_connect_retry=on`. Set @@ -2672,9 +2981,9 @@ you also set `query_pool_min=0` or enable query retries (see [Connection events](#connection-events)). Replay after a reconnect is at least once: a batch that QuestDB committed just -before the connection dropped is sent again. - - +before the connection dropped is sent again. Write to a deduplicated table, as +described under [Store-and-forward](#store-and-forward), to keep replayed rows +from inserting duplicates. ### Query failover @@ -2684,7 +2993,7 @@ endpoint when there is one, and runs the query again from the start: | Key | Default | Purpose | |---|---|---| | `failover` | `on` | Set `off` to fail the query instead of retrying. | -| `failover_max_attempts` | `8` | Connection attempts per failure. | +| `failover_max_attempts` | `8` | Connection attempts per failure. Each attempt tries every endpoint in `addr`. | | `failover_backoff_initial_ms` | `50` | First retry delay. | | `failover_backoff_max_ms` | `1000` | Longest delay between retries. | | `failover_max_duration_ms` | `30000` | Time budget per failure. | @@ -2692,8 +3001,8 @@ endpoint when there is one, and runs the query again from the start: The attempt limit and the time budget apply together, and whichever is reached first ends the failover. When attempts fail fast, for example with connection refused while a server restarts, the 8 attempts and their backoff of 50 ms to -1 second, with jitter, take only about 2 to 4 seconds, long before the -30-second budget. To ride out a longer restart, raise `failover_max_attempts`, +1 second, with jitter, take only about 1 to 3 seconds, and at most about 4.5 +seconds, long before the 30-second budget. To ride out a longer restart, raise `failover_max_attempts`, or set `maxAttempts: 0` in a typed `egressSession.reconnect` object to remove the attempt limit and rely on the time budget alone; see [Typed reconnect policy](#typed-reconnect-policy). @@ -2722,7 +3031,10 @@ for each query on its own, so it also covers concurrent queries on separate leases: ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { + connectQwpNodeClient, + QwpReconnectExhaustedError, +} from "@questdb/nodejs-client"; const db = await connectQwpNodeClient( "ws::addr=db-a.example.com:9000,db-b.example.com:9000;", @@ -2742,6 +3054,10 @@ try { } await query.completion; console.log(`${rows.length} rows`); + } catch (error) { + if (!(error instanceof QwpReconnectExhaustedError)) throw error; + // Failover gave up. This lease stays failed: return it, retry later. + console.error("no endpoint could run the query:", error.message); } finally { await lease.close(); } @@ -2750,6 +3066,20 @@ try { } ``` +The reset needs the rows of the current attempt in one place you can discard. +When your code passes rows on as they arrive, for example by streaming them to +an HTTP response, it cannot take them back after a restart. Then choose one +of these instead: + +- Run such queries on a separate client with `failover=off`, so that a lost + connection fails the query instead of restarting it, and retry the whole + request. +- Keep each result small enough to buffer, for example by paging with + `WHERE timestamp < $1 ORDER BY timestamp DESC LIMIT n`, binding the oldest + timestamp of the previous page. +- Treat `batchSequence === 0n` after rows have left as an error, and abort the + downstream response instead of sending duplicates. + To be notified of a restart, set `egressSession.onReplayReset` in the second argument of `connectQwpNodeClient()`. It runs before the first replayed batch is delivered, and its event has `requestId`, `endpoint`, `previousEndpoint`, @@ -2791,7 +3121,10 @@ The object's fields and the connect-string keys they replace: | `onEvent` | None | None | `egressSession: { reconnect: false }` is the typed equivalent of -`failover=off`. +`failover=off`. Senders in background memory mode or with `sf_dir` retry +indefinitely: `maxAttempts` and `maxDurationMs` do not end their retries, so +the full example's `reconnect: { onEvent }` with `sf_dir` keeps retrying +through an outage of any length. The two directions treat the first connection differently: @@ -2858,19 +3191,28 @@ retried; see [Typed reconnect policy](#typed-reconnect-policy). | `attempt-failed` | One connection attempt failed. The client retries if the error is retryable and its budget allows; otherwise this is the last event before the failure is reported. | | `reconnected` | Reconnected to the same endpoint. | | `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | -| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. Only senders that retry indefinitely wait: store-and-forward senders after their first connection, and senders with `initial_connect_retry=async` or `lazy_connect=on`. | +| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. Only senders that retry indefinitely wait: store-and-forward senders after their first connection, and senders in background memory mode. | | `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | | `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | +The `QWP_RECONNECT_EVENT_KIND` constants name these kinds: `CONNECTED`, +`RECONNECTING`, `ATTEMPT_FAILED`, `RECONNECTED`, `FAILED_OVER`, +`DURABLE_ACK_UNAVAILABLE`, `DURABLE_ACK_PERSISTENT_FAILURE`, and +`PRIMARY_UNAVAILABLE`. + `reconnected` and `failed-over` are mutually exclusive: code that tracks the -current node must handle both. A query that reconnects runs again from its -first batch; see [Query failover](#query-failover) for resetting accumulated -rows. +current node must handle both. The client has no property that says whether it +is connected right now. To report it, for example in a health check, track the +latest event: after `reconnecting` the connection is down, and `connected`, +`reconnected`, or `failed-over` mean it is up. A query that reconnects runs +again from its first batch; see +[Query failover](#query-failover) for resetting accumulated rows. No event marks a terminal failure. When a sender stops retrying, because its reconnect budget ran out or the error cannot be retried, -`ingressSession.onError` receives the error with `terminal: true`, even while -the sender is idle. The sender's next `flush()`, auto-flushing `at()`, or +`ingressSession.onError` receives a `QwpIngressErrorEvent` with +`terminal: true`, even while the sender is idle. The event also has `error`, +`timestampMs`, and, for a server rejection, `senderError`. The sender's next `flush()`, auto-flushing `at()`, or `close()` then rejects with the same error, such as `QwpReconnectExhaustedError`. A query that cannot fail over rejects its iteration and `completion` instead. @@ -2880,24 +3222,6 @@ acknowledged, and durably acknowledged sequences, and `onError`, for session errors. `sender.metrics` returns a snapshot of the sender's counters, including `metrics.ingress` with the replay queue, reconnect, and notification counters. -## Concurrency - -Node.js runs your code on one thread, but async functions interleave at every -`await`: - -- **`QwpClient`** is safe to share across your whole application. -- **Senders** are not safe for concurrent producers. A row is built across - several calls, so an `await` between `table()` and `at()` lets another task - add columns to the same row. Give each producer its own sender, borrowed from - the pool, and size `sender_pool_max` to match. -- **Query leases** run one query at a time. Borrow one lease per concurrent - query; `query_pool_max` caps concurrent queries. -- **Worker threads** cannot share clients. Create one client per worker, and - give each worker its own `sender_id` when using store-and-forward. - -Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy -work out of them. - ## Configuration reference @@ -2910,7 +3234,7 @@ every key. The Node.js client's defaults and deviations: | `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | | `username`, `password`, `token` | none | Basic or bearer authentication. | | `tls_verify`, `tls_roots` | `on`, Node.js CA bundle | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | -| `connect_timeout`, `auth_timeout_ms` | `15000` | TCP/TLS connection and upgrade deadlines, in milliseconds. | +| `connect_timeout`, `auth_timeout_ms` | `15000` | DNS and TCP/TLS connection, and upgrade deadlines, in milliseconds. See [Connection timeouts](#connection-timeouts). | | `auto_flush` | `on` | Master switch for the three triggers. | | `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | | `auto_flush_interval` | `100` | Milliseconds. `0` disables. `off` is rejected. | @@ -2920,12 +3244,12 @@ every key. The Node.js client's defaults and deviations: | `request_durable_ack` | `off` | Enterprise. | | `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | | `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | -| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in memory mode, without `sf_dir` or background replay. `0` removes it. | +| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in default memory mode. `0` removes it. | | `max_frame_rejections`, `poison_min_escalation_window_millis` | `4`, `300000` | Poison-frame detector: rejections of one batch, and the minimum time they must span, before the sender stops. | | `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | | `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | | `sf_durability` | `memory` | `memory`, `periodic`, or `append`. | -| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](#store-and-forward), not a hard disk limit; memory queue cap without `sf_dir`. | +| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](#sf-capacity), not a hard disk limit; memory queue cap without `sf_dir`. | | `sf_max_segment_bytes` | `4m` with `sf_dir`, none without | Journal segment size, which also caps a batch. | | `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | | `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. | @@ -2935,7 +3259,7 @@ every key. The Node.js client's defaults and deviations: | `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | | `client_id` | `typescript/` | Sent to the server for diagnostics. | | `error_inbox_capacity`, `connection_listener_inbox_capacity` | `256`, `64` | Queues for rejection callbacks and connection events. | -| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client only. | +| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client. A standalone `Sender` also applies `lazy_connect`. | The [API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) @@ -2951,16 +3275,21 @@ places: | Area | Node.js behavior | |---|---| -| Outage budget | A sender without `sf_dir` or background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. See [Ingestion reconnect](#ingestion-reconnect). | +| Outage budget | A sender in default memory mode gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. Senders in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) or with `sf_dir` retry indefinitely. See [Ingestion reconnect](#ingestion-reconnect). | | `target` and `zone` | Also apply to ingestion. Set a query-only role with the typed `egress.target` option. See [Multiple endpoints](#multiple-endpoints). | -| Authentication rejected after a first connection | Senders with `sf_dir` or background replay keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | -| Durable acknowledgement unavailable | Background senders keep retrying from startup, and store-and-forward senders after their first connection, emitting `durable-ack-unavailable`. See [Durable acknowledgement](#durable-acknowledgement). | +| Authentication rejected after a first connection | Senders with `sf_dir` or in background memory mode keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | +| Durable acknowledgement unavailable | Senders in background memory mode keep retrying from startup, and store-and-forward senders after their first connection, emitting `durable-ack-unavailable`. See [Durable acknowledgement](#durable-acknowledgement). | | `sf_durability` | Also accepts `append`. | -| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Store-and-forward](#store-and-forward). | -| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Store-and-forward](#store-and-forward). | +| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Journal capacity](#sf-capacity). | +| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Lock recovery](#sf-lock-recovery). | | `max_lifetime_ms` | Closes idle connections above the pool minimum only. Connections at the minimum are not recycled. | -| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush. Size values take single-letter suffixes only. `tls_roots` must be PEM. `init_buf_size` and `max_buf_size` are rejected. | -| Defaults | `connect_timeout` is `15000`, `poison_min_escalation_window_millis` is `300000`, and `close_flush_timeout_millis` is `5000`. | +| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush or from sender creation. Size values take single-letter suffixes only. `tls_roots` must be PEM. `init_buf_size` and `max_buf_size` are rejected. | +| `connect_timeout` | Also covers DNS and the TLS handshake, and `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. See [Connection timeouts](#connection-timeouts). | +| `tls_roots` default | The CA certificates bundled with Node.js, not the operating system's trust store. See [TLS](#tls). | +| Defaults | `connect_timeout` is `15000` and `poison_min_escalation_window_millis` is `300000`. `close_flush_timeout_millis` is `5000`, as in the Rust, C, C++, Python, and Go clients; Java and .NET use `60000`. | +| Close after an ACK timeout | A standalone sender's `close()` rejects with `QwpSenderCloseTimeoutError` instead of logging a warning. The pooled client's `db.close()` resolves and reports the timeout, best-effort, to `ingressSession.onError`. See [Closing a sender](#closing-a-sender). | +| Error reports | Categories and policies are lowercase, hyphenated strings, such as `schema-mismatch` and `retriable-other`. See [Ingestion errors](#ingestion-errors). | +| Pool and query keys on a standalone `Sender` | The `Sender` logs a warning for the pool and query-only keys it ignores. It applies `client_id` and `lazy_connect`. | | `on_*_error` keys | Accepted but not applied. | ## Migration @@ -3064,7 +3393,7 @@ For ILP options, see the [`SenderOptions` reference](https://questdb.github.io/nodejs-questdb-client/classes/_questdb_nodejs-client.SenderOptions.html) and the [ILP overview](/docs/connect/compatibility/ilp/overview/). -## Full example: ingestion and querying with failover +## Full example: Ingestion and querying with failover A production-oriented pattern that ingests trades and queries recent prices, with TLS, a token, several hosts, error handling, and failover handling. Before @@ -3110,7 +3439,8 @@ function alertOperator(message: string) { function logConnection(event: QwpReconnectEvent) { if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { - console.info("questdb connection:", event.kind, String(event.endpoint ?? "")); + const endpoint = String(event.endpoint ?? ""); + console.info("questdb connection:", event.kind, endpoint); } } @@ -3145,8 +3475,12 @@ const db = await connectQwpNodeClient( egressSession: { queryTimeoutMs: 30_000, // Replaces any failover* keys; omitted fields use the defaults. - // maxAttempts 0 removes the 8-attempt limit, so failover lasts 30 s. - reconnect: { maxAttempts: 0, maxDurationMs: 30_000, onEvent: logConnection }, + // No 8-attempt limit (maxAttempts 0): failover can last up to 30 s. + reconnect: { + maxAttempts: 0, + maxDurationMs: 30_000, + onEvent: logConnection, + }, onReplayReset: (event) => console.warn("query restarts on", String(event.endpoint)), }, @@ -3192,7 +3526,9 @@ try { } finally { // After a terminal rejection, close() rejects with the same failure. // Log it so that it does not replace the error thrown above. - await sender.close().catch((error) => console.error("close failed:", error)); + await sender + .close() + .catch((error) => console.error("close failed:", error)); } // Querying: rows may not be visible yet, see "Read-after-write". diff --git a/documentation/connect/compatibility/pgwire/nodejs.md b/documentation/connect/compatibility/pgwire/nodejs.md index 4d02825bcf..891ff8a154 100644 --- a/documentation/connect/compatibility/pgwire/nodejs.md +++ b/documentation/connect/compatibility/pgwire/nodejs.md @@ -790,9 +790,8 @@ latestByQuery() QuestDB's support for the PostgreSQL Wire Protocol allows you to use standard JavaScript PostgreSQL clients for querying time-series data. Both `pg` and `postgres` clients offer good performance and features for working with QuestDB. -We recommend the `pg` client for querying. -For data ingestion, consider the QuestDB [Node.js client](/docs/connect/clients/nodejs/), which also streams -query results over QWP. +Among PGWire drivers, we recommend the `pg` client for querying. For data ingestion, and for streaming query +results without a PostgreSQL driver, use the QuestDB [Node.js client](/docs/connect/clients/nodejs/), which speaks QWP. Remember that QuestDB is optimized for time-series data, so make the most of its specialized time-series functions like `SAMPLE BY` and `LATEST ON` for efficient queries. diff --git a/documentation/connect/overview.md b/documentation/connect/overview.md index 46c2df520f..dffe2b3978 100644 --- a/documentation/connect/overview.md +++ b/documentation/connect/overview.md @@ -66,6 +66,11 @@ Highlights: - **Schema-flexible** — automatic table creation and on-the-fly column additions. +The throughput and latency figures are peaks. Actual rates depend on the +client, the hardware, and the row shape: a Node.js process, for example, +encodes rows on a single CPU core. Measure with your own client and data +before sizing an ingestion tier. + Pick a language: diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index e46db0e4e7..c5ae8c9fae 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -76,9 +76,8 @@ Why each line matters: It bounds only the **blocking** initial connect (`initial_connect_retry=on` / `sync`). Once a sender is running, the reconnect loop never consults it and retries a transport outage forever. Setting a large value here does nothing for -a running producer. The exception is a Node.js sender with neither `sf_dir` -nor background replay (`initial_connect_retry=async`, or pooled -`lazy_connect=on`), which applies it to every outage. See +a running producer. The exception is a Node.js sender in default memory +mode, which applies it to every outage. See [Reconnect and outage handling](#reconnect-and-outage-handling). ::: @@ -203,14 +202,14 @@ callers block up to `acquire_timeout_ms` then throw. | `sf_durability` | `memory` (also supports `periodic`; Node.js also `append`) | | `sf_sync_interval_millis` | `5000` (requires `sf_durability=periodic`) | | `sf_append_deadline_millis` | `30000` | -| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** (a Node.js sender without `sf_dir` or background replay: every outage) | +| `reconnect_max_duration_millis` | `300000` — bounds the **blocking initial connect only** (Node.js default memory mode: every outage) | | `reconnect_initial_backoff_millis` | `100` | | `reconnect_max_backoff_millis` | `5000` | -| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/Node.js) | -| `connect_timeout` | unset (Node.js: `15000`) — per-endpoint TCP connect bound, must be `> 0` | +| `close_flush_timeout_millis` | `60000` (Java/.NET) · `5000` (Rust/C/C++/Python/Go/Node.js) | +| `connect_timeout` | unset (Node.js: `15000`, also bounding DNS and TLS) — per-endpoint TCP connect bound, must be `> 0` | | `auth_timeout_ms` | `15000` | | `max_frame_rejections` | `4` | -| `poison_min_escalation_window_millis` | `5000` (Node.js: `300000`) | +| `poison_min_escalation_window_millis` | `5000` (Node.js: `300000`; Go: not supported, its window is `reconnect_max_duration_millis`) | ### Query client @@ -227,9 +226,8 @@ callers block up to `acquire_timeout_ms` then throw. There is no "retry forever" setting to look for on the reconnect keys — a running sender already does. `reconnect_max_duration_millis` applies only to a -blocking initial connect, except on a Node.js sender with neither `sf_dir` -nor background replay; see -[Reconnect and outage handling](#reconnect-and-outage-handling). +blocking initial connect, except on a Node.js sender in default memory mode; +see [Reconnect and outage handling](#reconnect-and-outage-handling). --- @@ -373,11 +371,11 @@ fail when the buffer reaches `sf_max_total_bytes` and exhausts `sf_append_deadline_millis` on `append()`. The Node.js client retries authentication rejections after the first -connection only in senders with `sf_dir` or background memory replay +connection only in senders with `sf_dir` or in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`); see [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). -For unsupported durable acknowledgement, its senders with -`initial_connect_retry=async` or `lazy_connect=on` keep retrying from startup +For unsupported durable acknowledgement, its senders in background memory +mode keep retrying from startup and emit `durable-ack-unavailable`, and store-and-forward senders do the same after their first successful connection. Monitor these [connection events](/docs/connect/clients/nodejs/#connection-events) and buffer @@ -409,10 +407,12 @@ has outlasted your configuration. :::note Alignment This is the behaviour of the Java reference client and the .NET client. Other -clients are aligned to it, except a Node.js sender with neither `sf_dir` nor -background replay (`initial_connect_retry=async`, or pooled -`lazy_connect=on`), which gives up after `reconnect_max_duration_millis`; see -the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). If you +clients are aligned to it, except a Node.js sender in default memory mode, +without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, which +gives up after `reconnect_max_duration_millis`; see the +[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect) and its +[other differences](/docs/connect/clients/nodejs/#differences-from-other-clients). +If you are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, and apply back-pressure to the producer rather than dropping data. @@ -429,7 +429,7 @@ and apply back-pressure to the producer rather than dropping data. | HTTP upgrade timeout / non-auth transport error | try next endpoint | | `421` with `X-QuestDB-Role: REPLICA` | role reject; try next endpoint | | `401` / `403` auth failure | never try later endpoints; **terminal** before the first successful connection, then client-specific: [Java and some Node.js senders retry](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | -| durable-ack requested but unsupported | terminal mismatch, except that Java senders retry after their first successful connection, and Node.js senders retry in background modes, or after their first connection with `sf_dir` ([details](/docs/connect/clients/nodejs/#durable-acknowledgement)) | +| durable-ack requested but unsupported | terminal mismatch, except that Java senders retry after their first successful connection, and Node.js senders retry from startup in background memory mode, or after their first connection with `sf_dir` ([details](/docs/connect/clients/nodejs/#durable-acknowledgement)) | | successful write upgrade | bind this endpoint | | all endpoints fail transport | throw / retry per initial/reconnect mode | | all endpoints role-reject as replicas | `QwpRoleMismatchException` | diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index df07504d14..297a2cb2e5 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1134,8 +1134,8 @@ creation, instead of from the first buffered row. Ingress senders use a reconnect loop regardless of whether store-and-forward is configured. The two storage modes share the same failover semantics, apart -from the Node.js memory-mode budget in the table below; they differ only in -where unacknowledged data lives: +from the Node.js default-memory-mode budget in the table below; they differ +only in where unacknowledged data lives: - **`sf_dir` set** (store-and-forward): segments are memory-mapped files under `sf_dir`. Unacknowledged data survives sender restarts and is replayed by @@ -1151,7 +1151,7 @@ section of the connect string reference: | Key | Default | Description | |----------------------------------|-----------|-------------------------------------------| -| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`). | +| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on` ([details](/docs/connect/clients/nodejs/#differences-from-other-clients)). | | `reconnect_initial_backoff_millis` | `100` | First post-failure sleep. | | `reconnect_max_backoff_millis` | `5000` | Cap on per-attempt sleep. | | `initial_connect_retry` | `off` | Retry on first connect (`on`, `sync`, `async`). | @@ -1168,7 +1168,7 @@ Key behaviors: - **Authentication rejection (`401`/`403`) never moves to another host.** It is terminal before a sender's first successful connection. After that, the Java client retries it indefinitely, the Node.js client does so for senders - with `sf_dir` or background memory replay (`initial_connect_retry=async` or + with `sf_dir` or in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`), and other clients stop; see [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). - **`421 + X-QuestDB-Role`** is a role reject: transient if the role is diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index 1844d5663c..0a2bb2db11 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -81,7 +81,9 @@ the host re-advertises a different zone. follow the primary regardless of geography. Ingress is currently zone-blind in both storage modes, so the `zone=` key is silently accepted on ingress connections and only takes effect on egress. The Node.js client is the -exception: it applies `zone=` and `target=` to ingress too. +exception: it applies `zone=` and `target=` to ingress too. Its other +deviations are listed under +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ### Selection priority @@ -121,10 +123,8 @@ The `target=` key controls which server role the client is willing to bind to: up to its predecessor's WAL — the client treats it as transient and retries the same host (with a fresh round, no exponential backoff) until it becomes a full `PRIMARY`. On an ingress sender this retry has no deadline; the producer -is bounded by buffer capacity rather than by elapsed time. The exception is a -Node.js sender with neither `sf_dir` nor background replay -(`initial_connect_retry=async`, or pooled `lazy_connect=on`), which stops after -`reconnect_max_duration_millis`. +is bounded by buffer capacity rather than by elapsed time, except for the +Node.js sender described under [Ingress (writes)](#ingress-writes). A `421 Misdirected Request` response **without** an `X-QuestDB-Role` header is treated as a generic transport error, not a role reject — the client walks @@ -151,9 +151,9 @@ length, and what bounds your tolerance is buffer capacity - Maximum backoff: `5 s` - Per-outage budget: **none**. `reconnect_max_duration_millis` bounds only the blocking sync initial connect, and the running loop never consults it. - The exception is a Node.js sender with neither `sf_dir` nor background - replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`), which - gives up after `reconnect_max_duration_millis` and fails with + The exception is a Node.js sender in default memory mode, without `sf_dir`, + `initial_connect_retry=async`, or `lazy_connect=on`, which gives up after + `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`; see the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps @@ -253,14 +253,17 @@ depends on the client and on when it arrives: | Client | Before a sender's first successful connection | After it | |---|---|---| | Java | Terminal | Retried indefinitely | -| Node.js | Terminal | Retried indefinitely by senders with `sf_dir` or background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`). Terminal for other senders | +| Node.js | Terminal | Retried indefinitely by senders with `sf_dir` or in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`). Terminal for other senders | | Rust, C, C++, Python, Go, .NET | Terminal | Terminal | Query connections and orphan drainers treat the rejection as terminal in every -client. A sender that retries keeps buffering until authentication succeeds -again, bounded by its buffer capacity, so a credential change on the cluster -does not stop the producer. Watch its connection events rather than waiting -for an error. See +client, except that a Java orphan drainer whose token comes from a token +provider retries for a bounded time before quarantining the slot. A sender +that retries keeps buffering until authentication succeeds again, bounded by +its buffer capacity, so a credential change on the cluster does not stop the +producer. Monitor such a sender: the Java client reports each rejection to the +sender's error handler as a retriable `SECURITY_ERROR`, and the Node.js client +emits an `attempt-failed` connection event for each failed attempt. See [Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors) for the Node.js rules. diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 8a76aafaf4..0d751ea8c0 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -18,7 +18,8 @@ first. `zone` and `target` are accepted everywhere but only take effect on egress: ingress parsers accept and ignore them, so one connect string can serve both directions. The Node.js client is the exception: it applies both keys to -ingress too. +ingress too; its other deviations are listed under +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). They are documented in full on the [connect-string reference](/docs/connect/clients/connect-string#failover-keys); the table below summarises the failover-relevant subset. @@ -27,8 +28,8 @@ the table below summarises the failover-relevant subset. |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | | `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only.** Which server role the query client accepts. Accepted and ignored on an ingress connect string. The [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints) instead applies it to ingress too, so set its query-side role through the typed `egress` option. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | -| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. The Node.js client bounds DNS and TCP/TLS separately with `connect_timeout` (15 s by default); other clients may use the OS default. On Node.js, `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. Lower it for faster upgrade failure detection; tune the connect timeout separately. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only** (the Node.js client also applies it to ingress). Which server role the query client accepts. Other clients accept and ignore it on an ingress connect string. On the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), set the query-side role through the typed `egress` option instead. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. `connect_timeout` bounds the TCP connect separately: most clients leave it unset by default and then use the OS timeout, while Node.js defaults it to 15 s and also bounds DNS and TLS with it. On Node.js, `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. Lower `auth_timeout_ms` for faster upgrade failure detection; tune the connect timeout separately. | `addr` syntax — both of these are equivalent and produce the same three-peer list: @@ -49,7 +50,7 @@ for the full list. The failover-relevant keys are: | Key | Type | Default | Notes | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 6bb17c787a..7a3ea586a8 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -19,9 +19,9 @@ arrive asynchronously. A network outage or a server restart leaves your producer code unaffected — the I/O thread quietly reconnects and replays what remains. In SF mode, even a crash of the sender process itself loses no unacked data: the next sender on the slot recovers it from disk and -replays it. The one exception is a Node.js sender with neither `sf_dir` nor -background replay, whose `flush()` waits for the reconnect during an outage; -see [Reconnect and replay](#reconnect-and-replay). +replays it. The one exception is a Node.js sender in default memory mode, +whose `flush()` waits for the reconnect during an outage; see +[Reconnect and replay](#reconnect-and-replay). ## Two modes @@ -35,16 +35,18 @@ SF runs in either of two modes selected by the connect string: | Unacked data if the sender crashes | Lost | Recovered and replayed on restart | | Unacked data if the sender's host reboots | Lost | Recovered, if the disk persists | | Tolerates transient network blips | Yes | Yes | -| Tolerates multi-minute server outages | Bounded by RAM cap | Bounded by disk cap | +| Tolerates multi-minute server outages | Bounded by RAM cap (Node.js default memory mode: also by `reconnect_max_duration_millis`) | Bounded by disk cap | | Recovers another sender's stale slot | n/a | Opt-in via `drain_orphans=on` | Both modes share the same reconnect loop, the same backoff and retry budgets, and the same on-the-wire behaviour. The only difference is -where unacked data lives. The Node.js client is the exception: a sender with -neither `sf_dir` nor background replay (`initial_connect_retry=async`, or -pooled `lazy_connect=on`) gives up after `reconnect_max_duration_millis` -(5 minutes by default), while its other senders retry indefinitely. See the -[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). +where unacked data lives. The Node.js client is the exception: it splits +memory mode in two. A sender in default memory mode gives up after +`reconnect_max_duration_millis` (5 minutes by default), while a sender in +background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) +retries indefinitely, as SF mode does. See the +[Node.js ingestion modes](/docs/connect/clients/nodejs/#flushing) and +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ## What "frame" means here @@ -94,8 +96,7 @@ Two consequences: committed, so replay is **at least once** and can insert duplicate rows. `wireSeq` is transport bookkeeping, not a deduplication key. For idempotent ingestion, use stable source IDs and timestamps with table-level - `DEDUP UPSERT KEYS`, as in the [Node.js store-and-forward - example](/docs/connect/clients/nodejs/#store-and-forward). + `DEDUP UPSERT KEYS`; see [Deduplication](/docs/concepts/deduplication/). ## Trim: how unacked data is reclaimed @@ -138,9 +139,10 @@ object store** (S3, Azure Blob, GCS, or NFS). response and rejects a connection without it, rather than waiting for ack frames it cannot receive. In most clients the rejection is terminal. The Java client retries it after a sender's first successful connection, so a - capability change on the cluster cannot stop the producer. Node.js background senders - (`initial_connect_retry=async`, or pooled `lazy_connect=on`) retry it from - startup, and Node.js store-and-forward senders after their first + capability change on the cluster cannot stop the producer. Node.js senders + in background memory mode (`initial_connect_retry=async` or + `lazy_connect=on`) retry it from startup, and Node.js store-and-forward + senders after their first connection, emitting `durable-ack-unavailable` events; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). A retrying sender keeps buffering, so monitor it. @@ -162,12 +164,11 @@ the reconnect loop documented in The producer is **not notified**: it keeps publishing into the substrate, subject to available capacity (see [Backpressure](#backpressure)). -On Node.js, only a sender with neither `sf_dir` nor background replay waits -for the reconnect in `flush()`, up to `reconnect_max_duration_millis`. -Background memory mode, enabled by -`initial_connect_retry=async` or pooled `lazy_connect=on`, keeps accepting +On Node.js, only a sender in default memory mode waits for the reconnect in +`flush()`, up to `reconnect_max_duration_millis`. Background memory mode, +enabled by `initial_connect_retry=async` or `lazy_connect=on`, keeps accepting batches into the memory replay queue until capacity is exhausted and retries -indefinitely. See the [three Node.js flushing modes](/docs/connect/clients/nodejs/#flushing). +indefinitely. See the [three Node.js ingestion modes](/docs/connect/clients/nodejs/#flushing). On every successful (re)connect: @@ -198,7 +199,7 @@ call throws a typed exception. On Node.js with `sf_dir`, this value is a journal size target rather than a hard disk limit. Transaction-closing batches and retained symbol dictionaries can exceed it, and other metadata needs additional space. Provision headroom -for each sender; see [Node.js journal capacity](/docs/connect/clients/nodejs/#store-and-forward). +for each sender; see [Node.js journal capacity](/docs/connect/clients/nodejs/#sf-capacity). Without `sf_dir`, the key caps the in-memory replay queue. The exception message distinguishes the two scenarios: @@ -217,8 +218,8 @@ The exception message distinguishes the two scenarios: `close()` waits up to `close_flush_timeout_millis` for `ackedFsn` to reach `publishedFsn` — i.e. for the server to acknowledge everything the producer has handed in. The default differs by client: 60 s on Java and .NET, 5 s on Rust, -C, C++, Python and Node.js. If the wait succeeds, all data is acked. If the timeout -fires, a `WARN` is logged and: +C, C++, Python, Go and Node.js. If the wait succeeds, all data is acked. If the +timeout fires, a `WARN` is logged and: - in **SF mode**, the un-acked tail is left on disk and recovered by the next sender on the same slot; @@ -354,21 +355,24 @@ for the shared vocabulary and which clients let you override the defaults. | `WRITE_ERROR` | `retriable` | A write failed, for example under temporary storage pressure; reconnect and replay. | | `INTERNAL_ERROR` | `retriable` | Retry after an unexpected server-side failure. | | `DICTIONARY_GAP` | `retriable` | The connection is missing symbol dictionary entries; resend them and replay. | -| `NOT_WRITABLE` | `retriable_other` | The node cannot accept writes, for example a replica; replay on another endpoint. | +| `NOT_WRITABLE` | `retriable_other` | The node cannot accept writes, for example a replica; replay on another endpoint. Reserved: current servers close the connection instead, and the client reconnects. | | `UNKNOWN` | `retriable` (forced) | A status the client does not know, for example from a newer server; retry rather than stop. | A batch that keeps being rejected without progress escalates to a terminal error through the poison-frame detector (`max_frame_rejections`). The Java and Node.js clients also report a client-side `DATA_LOSS` category, with the `abandoned` policy, when they set aside a corrupt store-and-forward journal. +The Node.js client also defines `cancelled` and `limit-exceeded` categories, +both retriable; current servers send those statuses only to query +connections. Rejections are delivered asynchronously through a bounded error inbox (`error_inbox_capacity`, default `256`) that drops the oldest notification on overflow, to the application's error handler, such as `onSenderError` on Node.js. The default handler logs every rejection, because a silent handler -would hide data loss. The Node.js client reports categories as lowercase, -hyphenated `error.category` strings, such as `schema-mismatch`, and accepts -the `on_*_error` keys without applying them. +would hide data loss. The Node.js client reports categories and policies as +lowercase, hyphenated strings, such as `schema-mismatch` and +`retriable-other`, and accepts the `on_*_error` keys without applying them. ## Next steps diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index f9cf444ded..4a1c3158cf 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -27,7 +27,7 @@ mode. | `sf_dir` | path | unset | Group root directory. When set, the slot lives at `//` and unacked data is durable across process restarts. When unset, the substrate runs in memory mode. | | `sender_id` | string | `default` | Slot subdirectory name. Two senders sharing the same `sender_id` and `sf_dir` will collide on the slot lock. Must not contain path separators or be empty. | | `sf_max_segment_bytes` | size | `4M` | Per-segment file size; rotation threshold. | -| `sf_max_total_bytes` | size | `128M` (memory) / `10G` (SF) | Capacity for producer backpressure. Node.js disk journals can exceed this target for transaction completion and symbol dictionaries; provision [additional disk headroom](/docs/connect/clients/nodejs/#store-and-forward). Without `sf_dir`, this caps the memory replay queue. | +| `sf_max_total_bytes` | size | `128M` (memory) / `10G` (SF) | Capacity for producer backpressure. Node.js disk journals can exceed this target for transaction completion and symbol dictionaries; provision [additional disk headroom](/docs/connect/clients/nodejs/#sf-capacity). Without `sf_dir`, this caps the memory replay queue. | | `sf_durability` | enum | `memory` | `memory` relies on the page cache; `periodic` checkpoints in the background and requires `sf_dir`. Node.js also supports `append`, which makes each journal append durable before `flush()` resolves. Go and .NET accept only `memory`; other clients reject `append` at build time. `flush` is not supported. | | `sf_sync_interval_millis` | int (ms) | `5000` | Checkpoint cadence for `sf_durability=periodic`; rejected without it. A floor, not a guarantee: scheduler and storage latency add to it. | | `sf_append_deadline_millis` | int (ms) | `30000` | How long a producer `appendBlocking` call waits for ACK-driven trim to free space before throwing. | @@ -48,11 +48,11 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender with neither `sf_dir` nor background replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | -| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python and Node.js | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | +| `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python, Go and Node.js | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | Cross-reference: [connect-string #reconnect-keys](/docs/connect/clients/connect-string#reconnect-keys). @@ -72,13 +72,14 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| | `error_inbox_capacity` | int (≥16) | `256` | Bounded SPSC queue capacity for async error notifications. Overflow drops the oldest entry and increments `getDroppedErrorNotifications`. | -| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | All clients accept these keys. Go and .NET apply them; Java, Node.js, Rust, C, C++, and Python currently ignore them. There is no `DROP_AND_CONTINUE` policy. See [Error handling](/docs/connect/clients/connect-string/#error-handling). | +| `on_server_error`, `on_schema_error`, `on_parse_error`, `on_internal_error`, `on_security_error`, `on_write_error` | enum | per category | All clients accept these keys. Go and .NET apply them; Java, Node.js, Rust, C, C++, and Python currently ignore them. The two that apply them accept different values: .NET takes only `halt`/`terminal` and `retry`/`retriable` and rejects `auto` and `retriable_other`. There is no `DROP_AND_CONTINUE` policy. See [Error handling](/docs/connect/clients/connect-string/#error-handling). | The per-category defaults are documented in [Concepts § Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). -`PROTOCOL_VIOLATION` is always terminal and `UNKNOWN` always retriable, so a -status from a newer server leads to a retry rather than a silently dropped -batch. +`PROTOCOL_VIOLATION` is always terminal and `UNKNOWN` defaults to retriable, +so a status from a newer server leads to a retry rather than a silently +dropped batch. The `on_*_error` keys cannot change either; only the .NET +client's programmatic resolver can change `UNKNOWN`. ## Other relevant keys diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index f32ab071ca..bbf673e3e8 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -62,9 +62,9 @@ keeps `.lock` and `.lock.pid` only for compatibility. After a crash, a new Node.js sender takes the slot over automatically only on the same host, once the recorded process ID is no longer in use. In containers that usually fails, because the application runs as process ID 1 and a replacement container has a -new host name, and the new sender reports `QwpReplayStoreLockedError`. The -[Node.js client](/docs/connect/clients/nodejs/#store-and-forward) describes how -to remove a stale lock safely. +new host name, and the new sender reports `QwpReplayStoreLockedError`. +[Node.js lock recovery](#nodejs-lock-recovery) describes how to remove a stale +lock safely. Node.js and other clients do not see each other's locks, so never let them use the same `sf_dir` at the same time. @@ -110,8 +110,8 @@ when the new one comes up. Solutions: - Stop the previous process. For clients using OS locks, the kernel releases the lock on exit (even after `kill -9`). A killed Node.js sender can leave a stale `.lock.owner` directory behind: verify that the old owner is gone - before removing it, as the - [Node.js client](/docs/connect/clients/nodejs/#store-and-forward) describes. + before removing it, as described under + [Node.js lock recovery](#nodejs-lock-recovery). - Use a deployment unit that orders shutdown before startup. - For containerised deployments, set `sender_id` from a per-pod stable identity so two pods with the same template name don't collide. @@ -119,6 +119,33 @@ when the new one comes up. Solutions: `drain_orphans=on` does **not** override the lock — a busy orphan slot is skipped, not stolen. +### Node.js lock recovery {#nodejs-lock-recovery} + +A Node.js sender that crashed, or ran in a container that was replaced, can +leave its `.lock.owner` directory behind, and a new sender on the slot then +fails with `QwpReplayStoreLockedError`. To recover: + +1. Verify that the previous owner has exited and that no process is using the + slot. The `.lock.owner` directory records the owner's host name and process + ID. +2. Remove the stale `//.lock.owner` directory, where `` is + ``, or `-` for a pooled sender, and start the + client again. +3. If startup still reports `QwpReplayStoreLockedError`, also inspect + `/.slot-locks/.lock.owner`. This short-lived guard can + survive a crash during lock acquisition or quarantine. Remove that specific + owner directory only after verifying that its owner has exited. + +Never delete the shared `.slot-locks` directory or another slot's locks. + +Automate this cleanup only where the deployment guarantees that the previous +owner has exited before a new one starts, for example a single replica that +uses the Kubernetes `Recreate` update strategy and a `ReadWriteOnce` volume. A +startup step can then remove the stale owner directories of the client's own +slots before it creates the client. Anywhere two processes can overlap, +recover manually. See also the +[Node.js client](/docs/connect/clients/nodejs/#sf-lock-recovery). + ## Sizing capacity Two limits matter: @@ -150,7 +177,7 @@ commit possible when the journal is full. Segment reservations can reach roughly twice the target, depending on segment rounding, with retained symbol dictionaries and other metadata requiring additional space. Provision headroom per sender and monitor actual disk usage; do not use the target as a -filesystem quota. See [Node.js store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). +filesystem quota. See [Node.js journal capacity](/docs/connect/clients/nodejs/#sf-capacity). ::: @@ -216,7 +243,7 @@ fresh start: no segments, no replay. | Symptom | Likely cause | Operator action | |---|---|---| -| "Slot held by PID ``" or `QwpReplayStoreLockedError` (Node.js) | Another process holds the slot, or a Node.js `.lock.owner` is stale after a crash. | Stop the duplicate. OS locks release on exit; for Node.js verify the owner is gone before removing `.lock.owner` (see the [Node.js client](/docs/connect/clients/nodejs/#store-and-forward)). | +| "Slot held by PID ``" or `QwpReplayStoreLockedError` (Node.js) | Another process holds the slot, or a Node.js `.lock.owner` is stale after a crash. | Stop the duplicate. OS locks release on exit; for Node.js verify the owner is gone before removing `.lock.owner` (see [Node.js lock recovery](#nodejs-lock-recovery)). | | "Gap between segments" | Corruption — a segment was deleted out of band. | Restore from backup or accept data loss; the substrate refuses to start. | | "Watermark exceeds publishedFsn" | `.ack-watermark` is corrupt; the engine falls back to the no-watermark seed. | Logged as `WARN`. Replay will re-send the lowest segment's frames, which inserts duplicate rows unless the table has `DEDUP UPSERT KEYS`. | | Torn tail count > 0 | The previous process crashed mid-frame-write. | Informational; the CRC + zero-fill design discards the partial frame. | @@ -227,7 +254,7 @@ fresh start: no segments, no replay. | Value | Behaviour | |---|---| -| client default: `60000` on Java and .NET, `5000` on Rust, C, C++, Python, and Node.js | Block up to that long waiting for `ackedFsn ≥ publishedFsn`. Log `WARN` on timeout; un-acked tail stays on disk (SF) or is lost (memory). | +| client default: `60000` on Java and .NET, `5000` on Rust, C, C++, Python, Go, and Node.js | Block up to that long waiting for `ackedFsn ≥ publishedFsn`. Log `WARN` on timeout; un-acked tail stays on disk (SF) or is lost (memory). | | `0` or `-1` | Skip the drain wait. Pending data persists on disk (SF) for the next sender, or is lost (memory). | | any other positive value | That timeout in milliseconds. | diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 1d1f35cc44..7d862e7f40 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -57,9 +57,10 @@ Unacked frames are written to mmap'd files under Both modes share the same wire behaviour, the same failover loop, and the same connect-string keys for everything other than storage. You can switch between them without changing application code — only the connect -string. On the Node.js client, a sender with neither `sf_dir` nor background -replay (`initial_connect_retry=async`, or pooled `lazy_connect=on`) also gives -up after `reconnect_max_duration_millis` of outage. +string. On the Node.js client, a sender in default memory mode, without +`sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, also gives up +after `reconnect_max_duration_millis` of outage; see +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ## Comparison at a glance @@ -113,8 +114,8 @@ GCS, or NFS). or an uninitialised primary, the connection attempt is rejected. In most clients this is terminal. The exceptions follow. - **Senders that keep retrying.** The Java client retries after a sender's - first successful connection. Node.js senders with - `initial_connect_retry=async` or `lazy_connect=on` retry from startup + first successful connection. Node.js senders in background memory mode + (`initial_connect_retry=async` or `lazy_connect=on`) retry from startup instead of failing initialization, and Node.js store-and-forward senders retry after their first successful connection; they emit `durable-ack-unavailable` diff --git a/package.json b/package.json index b9d5f87499..82d1eff578 100644 --- a/package.json +++ b/package.json @@ -9,7 +9,8 @@ "build": "cross-env NO_UPDATE_NOTIFIER=true USE_SIMPLE_CSS_MINIFIER=true PWA_SW_CUSTOM= docusaurus build", "deploy": "docusaurus deploy", "serve": "docusaurus serve", - "swizzle": "docusaurus swizzle" + "swizzle": "docusaurus swizzle", + "test": "node --test \"plugins/**/*.test.js\"" }, "dependencies": { "@docusaurus/faster": "^3.8.1", diff --git a/plugins/raw-markdown/convert-components.js b/plugins/raw-markdown/convert-components.js index 29723cc450..172b065c9a 100644 --- a/plugins/raw-markdown/convert-components.js +++ b/plugins/raw-markdown/convert-components.js @@ -627,9 +627,11 @@ function removeImports(content) { for (let i = 0; i < lines.length; i++) { const line = lines[i] - const fence = line.match(/^ {0,3}(`{3,}|~{3,})(.*)$/) + // Ignore a CRLF line ending, so fences in CRLF files are still recognized. + const fence = line.replace(/\r$/, '').match(/^ {0,3}(`{3,}|~{3,})(.*)$/) if (!fenceChar) { - if (!fence) continue + // CommonMark: a backtick fence's info string cannot contain a backtick. + if (!fence || (fence[1][0] === '`' && fence[2].includes('`'))) continue appendOutside(i) fenceChar = fence[1][0] fenceLen = fence[1].length diff --git a/plugins/raw-markdown/convert-components.test.js b/plugins/raw-markdown/convert-components.test.js index f3553e807a..adf9ab35ca 100644 --- a/plugins/raw-markdown/convert-components.test.js +++ b/plugins/raw-markdown/convert-components.test.js @@ -31,6 +31,29 @@ test('removes MDX imports but keeps TypeScript imports in fenced examples', () = assert.match(result, /~~~ts\nimport \{ client \} from "@questdb\/nodejs-client";\n~~~/) }) +test('recognizes fences in files with CRLF line endings', () => { + const markdown = [ + 'import Widget from "@site/src/components/Widget"', + '```typescript', + 'import { Sender } from "@questdb/nodejs-client";', + '```', + ].join('\r\n') + + const result = removeImports(markdown) + assert.doesNotMatch(result, /@site\/src\/components\/Widget/) + assert.match(result, /import \{ Sender \} from "@questdb\/nodejs-client";/) +}) + +test('does not open a fence on backticks whose info string has a backtick', () => { + const markdown = [ + '```not `a fence`', + 'import Widget from "@site/src/components/Widget"', + ].join('\n') + + const result = removeImports(markdown) + assert.doesNotMatch(result, /@site\/src\/components\/Widget/) +}) + test('preserves imports from the Node.js Quick start while removing its MDX import', () => { const file = path.join(__dirname, '../../documentation/connect/clients/nodejs.md') const { content } = matter(fs.readFileSync(file, 'utf8')) diff --git a/shared/clients.json b/shared/clients.json index 0bec77694f..393030b9df 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -112,7 +112,7 @@ "protocol": "PGWire" }, { - "href": "/docs/connect/clients/nodejs", + "href": "/docs/connect/clients/nodejs#ilp-transports-legacy", "name": "Node.js", "description": "Node.js client for ILP ingestion over HTTP and TCP, alongside QWP.", From 353084dc325243f5888ecaf9e69f00d77bec5435 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 15:29:30 +0100 Subject: [PATCH 16/25] docs(nodejs): fix QWP examples and delivery guidance --- documentation/concepts/delivery-semantics.md | 7 +- documentation/connect/clients/nodejs.md | 154 ++++++++++++------ .../partials/_sf-dedup-warning.partial.mdx | 12 +- documentation/query/overview.md | 8 +- 4 files changed, 122 insertions(+), 59 deletions(-) diff --git a/documentation/concepts/delivery-semantics.md b/documentation/concepts/delivery-semantics.md index 0f7a791ad2..201651a5a4 100644 --- a/documentation/concepts/delivery-semantics.md +++ b/documentation/concepts/delivery-semantics.md @@ -121,12 +121,13 @@ If two distinct events can share `(ts, symbol, side)` and both should be preserved, widen `UPSERT KEYS` to include a column that distinguishes them — for example a `trade_id` or `seq` column. -:::warning DEDUP is required on tables behind multi-host failover +:::warning Use DEDUP behind multi-host failover when duplicates matter When the client fails over from one primary to another, unacknowledged batches are replayed against the new primary. Without `DEDUP UPSERT KEYS` -covering row identity, those replays produce duplicate rows in the target -table. +covering row identity, those replays can produce duplicate rows in the target +table. Enable DEDUP for exactly-once outcomes; applications that tolerate +occasional duplicates can skip it. ::: diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 524708276e..02393e126f 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -66,12 +66,13 @@ existing ILP code to QWP, see [From ILP to QWP](#from-ilp-to-qwp). The ## Installation ```shell -npm install @questdb/nodejs-client +npm install @questdb/nodejs-client@^5 ``` -The package also installs with `yarn add` and `pnpm add`. It exports its -complete API from the package root, ships ES module and CommonJS builds, and -bundles TypeScript declarations. There are no other supported import paths. +Use `yarn add @questdb/nodejs-client@^5` or +`pnpm add @questdb/nodejs-client@^5` with the other package managers. The +package exports its complete API from the package root, ships ES module and +CommonJS builds, and bundles TypeScript declarations. There are no other supported import paths. The examples on this page are TypeScript ES modules with top-level `await`. To run the [quick start](#quick-start) as TypeScript, save its code as @@ -91,7 +92,8 @@ TypeScript assertions such as `as const`. ## Quick start Connect with one connect string, create a table, write two rows, wait until -QuestDB acknowledges them, and query them back. +QuestDB acknowledges them, and run a query. QuestDB applies acknowledged rows +asynchronously, so the first query may return no rows. ```typescript import { @@ -536,8 +538,10 @@ try { .doubleColumn("amount", 0.1) .at(Date.now(), "ms"); } + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence); } finally { - // Flushes and returns the sender to the pool. Does not wait for ACKs. + // Returns the sender to the pool; any pending rows are flushed. await sender.close(); } } finally { @@ -550,8 +554,9 @@ A long-running producer can keep its borrow for its whole lifetime and call that hold a sender at the same time. `close()` on a borrowed sender flushes its completed rows and returns it to -the pool. It does not wait for acknowledgements, but in the default memory -mode its flush waits for the reconnect during an outage. See +the pool. The example above waits for the acknowledgement *before* returning +the sender: `close()` itself does not wait for acknowledgements, although in +default memory mode its flush waits for the reconnect during an outage. See [Closing a borrowed sender](#closing-a-borrowed-sender) for how long that can take, how to wait for acknowledgements, and what happens when a close fails. @@ -682,7 +687,10 @@ buffer rows in memory until QuestDB is reachable. The query pool stays empty until the first query. ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { + connectQwpNodeClient, + QwpIngressAckTimeoutError, +} from "@questdb/nodejs-client"; // Resolves immediately, even if QuestDB is not running yet. const db = await connectQwpNodeClient( @@ -698,6 +706,18 @@ try { .doubleColumn("price", 2615.54) .doubleColumn("amount", 0.5) .at(Date.now(), "ms"); + await sender.flush(); + const sequence = sender.publishedSequence; + // Keep the process running until QuestDB comes back and acknowledges it. + for (;;) { + try { + await sender.waitForAcknowledged(sequence, 10_000); + break; + } catch (error) { + if (!(error instanceof QwpIngressAckTimeoutError)) throw error; + console.info("still waiting for QuestDB; do not restage the row"); + } + } } finally { await sender.close(); } @@ -706,9 +726,10 @@ try { } ``` -Rows buffered while QuestDB is down exist only in memory. They are lost if the -client closes before QuestDB becomes reachable, as `db.close()` does at the end -of this example; see [Closing the pooled client](#closing-the-pooled-client). +Rows buffered while QuestDB is down exist only in memory. This example keeps +the client running until the row is acknowledged; if the process exits first, +the unacknowledged row may be lost. See +[Closing the pooled client](#closing-the-pooled-client). To keep them across a shutdown or restart, add a [store-and-forward](#store-and-forward) journal with `sf_dir`. Replay from the journal is at least once, so write to a deduplicated table as described there. @@ -1161,9 +1182,9 @@ try { ``` - `decimalColumnText()` takes a - decimal string, scientific notation included (`"1.5e-3"`), and preserves the - literal's scale, including trailing zeros. Passing a `number` works, but - JavaScript drops trailing zeros when formatting. + plain decimal string (such as `"0.0750"`) and preserves the literal's scale, + including trailing zeros. Scientific notation is accepted for a `number`, + not a string; JavaScript drops trailing zeros when formatting numbers. - `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` or as big-endian two's-complement bytes in an `Int8Array`. @@ -1703,7 +1724,8 @@ import { const db = await connectQwpNodeClient( "ws::addr=localhost:9000;" + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + - "sf_durability=append;lazy_connect=on;", + // Offline batches must fit the default server limit (about 2 MiB). + "sf_durability=append;sf_max_segment_bytes=1m;lazy_connect=on;", ).catch((error: unknown) => { // Another process holds the journal, or a crash left a stale lock. if ( @@ -2012,7 +2034,10 @@ Batch objects stay valid after iteration moves on, so you can keep them. With the default `failover=on`, a lost connection can make the client run the query again from its first batch. If your loop accumulates rows, reset them -when `batch.batchSequence === 0n`; see [Query failover](#query-failover). +when `batch.batchSequence === 0n`. A replay that returns **no batches** has no +sequence to detect, so also clear accumulated state on `onReplayReset` (on a +client with only one active query), or use `failover=off` and retry the whole +query. See [Query failover](#query-failover). ### Reading result values @@ -2254,7 +2279,18 @@ leave the poll running indefinitely: import { randomUUID } from "node:crypto"; import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +let visible = false; +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;query_pool_max=1;", + // There is only one active query. Clear its state even if replay returns no rows. + { + egressSession: { + onReplayReset: () => { + visible = false; + }, + }, + }, +); try { const lease = await db.borrowQuery(); try { @@ -2274,12 +2310,12 @@ try { .symbol("symbol", "ETH-USD") .at(Date.now(), "ms"); await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence); } finally { await sender.close(); } const deadline = Date.now() + 10_000; - let visible = false; while (!visible) { const remainingMs = deadline - Date.now(); if (remainingMs <= 0) throw new Error("trade not visible in time"); @@ -2311,8 +2347,12 @@ try { ``` Because the table already exists, an SQL error fails the loop at once instead -of being retried; the loop retries only while the row is not yet visible. Do -not replace the poll with a fixed sleep: the apply latency varies with load. +of being retried; the loop retries only while the row is not yet visible. The +single-query pool lets `onReplayReset` reset `visible` even when a replay +returns no batches. For concurrent queries, use separate clients for this +pattern, or set `failover=off` and retry the entire poll after a connection +error. Do not replace the poll with a fixed sleep: apply latency varies with +load. ### Cancellation and timeouts @@ -2480,12 +2520,13 @@ what your loop must do after `cancel()`. `query()` materializes every value into JavaScript arrays. For hot paths, `queryViews()` hands a reusable view of each batch to a callback, and reads -values straight from the received bytes: +values straight from the received bytes. This single-pass sum disables query +failover so a transport error rejects instead of leaving a partial sum: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); try { const lease = await db.borrowQuery(); try { @@ -2493,8 +2534,6 @@ try { const query = await lease.queryViews( "SELECT timestamp, symbol, price, amount FROM trades", (batch) => { - // A failover replays from batch 0; discard the failed attempt's sum. - if (batch.batchSequence === 0n) notional = 0; const price = batch.column(2); const amount = batch.column(3); for (let row = 0; row < batch.rowCount; row++) { @@ -2502,8 +2541,6 @@ try { notional += price.getDouble(row) * amount.getDouble(row); } } - // Row-major access reuses one row object for every row. - batch.forEachRow((r) => void r.getSymbol(1)); }, ); await query.completion; @@ -2516,6 +2553,11 @@ try { } ``` +For row-major work instead, `batch.forEachRow()` reuses one row object; read +values inside its callback only when they contribute to your result. For an +accumulator that supports automatic query replay, see +[Query failover](#query-failover). + Batches are delivered one at a time: when the callback returns a promise, the client waits for it before delivering the next batch. With a credit window set (`initialCredit`), it also grants credit for a batch only after its callback @@ -3025,10 +3067,12 @@ the query restarts; otherwise it sees the first part of the result twice. ::: -Detect the restart inside the loop. Every batch has a `batchSequence` that -starts at `0n`, and a re-executed query starts again at `0n`. The check works -for each query on its own, so it also covers concurrent queries on separate -leases: +Every batch has a `batchSequence` starting at `0n`, including the first batch +after a replay. Clear accumulated rows on that batch. A replay may return +**zero rows and no batches**, though, leaving prior rows in your accumulator. +`egressSession.onReplayReset` also clears it when that happens. Since this +callback is shared across the pool and request IDs are per connection, the +example limits the pool to one active query: ```typescript import { @@ -3036,8 +3080,16 @@ import { QwpReconnectExhaustedError, } from "@questdb/nodejs-client"; +const rows: (readonly unknown[])[] = []; const db = await connectQwpNodeClient( - "ws::addr=db-a.example.com:9000,db-b.example.com:9000;", + "ws::addr=db-a.example.com:9000,db-b.example.com:9000;query_pool_max=1;", + { + egressSession: { + onReplayReset: () => { + rows.length = 0; + }, + }, + }, ); try { const lease = await db.borrowQuery(); @@ -3046,7 +3098,6 @@ try { // The deadline covers the whole query, including a re-execution. timeoutMs: 30_000, }); - const rows: (readonly unknown[])[] = []; for await (const batch of query) { // Sequence 0 starts the result, both initially and after a failover. if (batch.batchSequence === 0n) rows.length = 0; @@ -3080,13 +3131,15 @@ of these instead: - Treat `batchSequence === 0n` after rows have left as an error, and abort the downstream response instead of sending duplicates. -To be notified of a restart, set `egressSession.onReplayReset` in the second -argument of `connectQwpNodeClient()`. It runs before the first replayed batch -is delivered, and its event has `requestId`, `endpoint`, `previousEndpoint`, -`serverInfo`, and `cause`. The `requestId` matches `query.requestId`, but -request IDs are numbered per connection and every lease of a pooled client -shares the callback, so the event cannot tell concurrent queries apart. Use it -for logging, and the sequence check above to reset results. +`egressSession.onReplayReset` runs before a query is replayed, including when +that replay returns no batches. Its event has `requestId`, `endpoint`, +`previousEndpoint`, `serverInfo`, and `cause`. The `requestId` matches +`query.requestId`, but request IDs are numbered per connection and every lease +of a pooled client shares the callback: it cannot identify which of several +concurrent queries restarted. Use a dedicated, single-query client when the +callback resets result state, as above. For concurrent results that cannot be +isolated, set `failover=off` and retry the whole query after a transport error; +`batchSequence === 0n` alone cannot detect a zero-batch replay. ### Typed reconnect policy @@ -3432,6 +3485,9 @@ import { const token = process.env.QDB_TOKEN; if (!token) throw new Error("QDB_TOKEN is not set"); +// This example runs one query at a time so its replay callback can clear it. +const recentPrices: (readonly unknown[])[] = []; + // Replace with your alerting. function alertOperator(message: string) { console.error("ALERT:", message); @@ -3449,9 +3505,11 @@ const db = await connectQwpNodeClient( `token=${token};` + // append: every flush waits for a disk sync; see "Store-and-forward". "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + - "sf_durability=append;sender_pool_max=4;" + - // query_pool_min=0: start, and ingest, even while no replica is reachable. - "query_pool_min=0;query_pool_max=8;", + // Limit offline batches below the default server's 2 MiB limit. + "sf_durability=append;sf_max_segment_bytes=1m;sender_pool_max=4;" + + // Query pool stays cold while replicas are down; one active query so the + // replay callback below can reset its state even if replay has no batches. + "query_pool_min=0;query_pool_max=1;", { // Queries run on replicas only, never on the primary; ingestion always // follows the primary. @@ -3481,8 +3539,10 @@ const db = await connectQwpNodeClient( maxDurationMs: 30_000, onEvent: logConnection, }, - onReplayReset: (event) => - console.warn("query restarts on", String(event.endpoint)), + onReplayReset: (event) => { + recentPrices.length = 0; + console.warn("query restarts on", String(event.endpoint)); + }, }, }, ); @@ -3540,9 +3600,9 @@ try { "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, ); - const recentPrices: (readonly unknown[])[] = []; for await (const batch of query) { - // A failover restarts the result at sequence 0: drop partial rows. + // A nonempty replay starts at sequence 0; the callback handles + // empty ones. if (batch.batchSequence === 0n) recentPrices.length = 0; for (const row of batch.rows()) recentPrices.push(row); } diff --git a/documentation/partials/_sf-dedup-warning.partial.mdx b/documentation/partials/_sf-dedup-warning.partial.mdx index d01c79f1f4..35d4193918 100644 --- a/documentation/partials/_sf-dedup-warning.partial.mdx +++ b/documentation/partials/_sf-dedup-warning.partial.mdx @@ -1,11 +1,11 @@ -:::caution Replay is at-least-once — enable DEDUP +:::caution Replay is at-least-once — use DEDUP for exactly-once outcomes After a reconnect or a sender restart, the client replays frames the server may have accepted but not yet acknowledged. Without -[DEDUP](/docs/concepts/deduplication/) on the target table, replay produces -duplicate rows. Tables ingested over a reconnecting or multi-host connection -**must** declare `DEDUP UPSERT KEYS(...)` covering row identity. See -[Delivery semantics](/docs/concepts/delivery-semantics/) for the full -at-least-once / exactly-once model. +[DEDUP](/docs/concepts/deduplication/) on the target table, replay can produce +duplicate rows. If your application requires exactly-once outcomes, declare +`DEDUP UPSERT KEYS(...)` covering row identity. Applications that tolerate +occasional duplicates can skip DEDUP; see +[Delivery semantics](/docs/concepts/delivery-semantics/) for the full model. ::: diff --git a/documentation/query/overview.md b/documentation/query/overview.md index c422f27212..9167063fb5 100644 --- a/documentation/query/overview.md +++ b/documentation/query/overview.md @@ -102,9 +102,11 @@ string. Ingestion and queries run over separate WebSocket connections, which the clients' pools manage for you. Results stream rather than arriving in one block. The server sends batches as -it produces them, so an application starts processing the head of a result -while the tail is still being computed, and a result larger than memory never -has to be materialized at all. +it produces them, so an application can process the head of a result while +the tail is still being computed without materializing the whole result. +Clients may buffer batches ahead of the consumer: for large Node.js results, +set a [byte-credit window](/docs/connect/clients/nodejs/#flow-control) to +bound client-side buffering (the default is unbounded). Connections can recover from transport failures. With failover enabled, a client may reconnect and re-execute an in-flight query, including on the same From 1fd38e55cdfd5f6eab3e5e3dd3e9329850a22b81 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 16:45:12 +0100 Subject: [PATCH 17/25] docs: correct Node.js QWP guidance and split operations reference --- documentation/concepts/delivery-semantics.md | 2 +- .../connect/clients/connect-string.md | 29 +- .../connect/clients/nodejs-operations.md | 1413 +++++++++++++++ documentation/connect/clients/nodejs.md | 1510 +---------------- .../wire-protocols/qwp-client-behavior.md | 19 +- .../wire-protocols/qwp-ingress-websocket.md | 2 +- .../client-failover/concepts.md | 6 +- .../client-failover/configuration.md | 8 +- .../store-and-forward/concepts.md | 15 +- .../store-and-forward/configuration.md | 2 +- .../store-and-forward/when-to-use.md | 14 +- documentation/query/overview.md | 6 +- documentation/sidebars.js | 5 + shared/clients.json | 2 +- 14 files changed, 1547 insertions(+), 1486 deletions(-) create mode 100644 documentation/connect/clients/nodejs-operations.md diff --git a/documentation/concepts/delivery-semantics.md b/documentation/concepts/delivery-semantics.md index 201651a5a4..ec6c3dc439 100644 --- a/documentation/concepts/delivery-semantics.md +++ b/documentation/concepts/delivery-semantics.md @@ -19,7 +19,7 @@ unacknowledged rows live in memory and are lost if the process exits, or the sender closes, before the server acknowledges them. A Node.js sender in default memory mode also gives up after `reconnect_max_duration_millis` of outage; see the -[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). +[Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). This page explains where duplicates come from and how to suppress them. diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 0e60f12f66..c92fdbce79 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -20,8 +20,8 @@ configures both without edits. The Node.js client is the exception for `target` and `zone`, which it also applies to ingress; see [Role filter and zone preference](#role-filter-and-zone-preference). The *Applies to:* tag on each section below marks which direction a key affects. -The [Node.js client page](/docs/connect/clients/nodejs/#differences-from-other-clients) -lists every key where that client differs from this reference. +The [Node.js client page](/docs/connect/clients/nodejs-operations/#differences-from-other-clients) +lists its behavioral differences from this reference. For legacy InfluxDB Line Protocol (ILP) transports (`http`, `https`, `tcp`, `tcps`), see the [ILP overview](/docs/connect/compatibility/ilp/overview/). @@ -454,7 +454,7 @@ The Node.js client applies `target` and `zone` to ingress as well. With `target=replica` in a shared connect string, its senders accept only replicas and cannot ingest. Set the query-side role through the typed `egress` option instead; see the -[Node.js client page](/docs/connect/clients/nodejs/#multiple-endpoints). +[Node.js client page](/docs/connect/clients/nodejs-operations/#multiple-endpoints). ::: @@ -487,15 +487,15 @@ server-side HA separately. Related: [Reconnect and failover](#reconnect-keys), [Store-and-forward](#sf-keys). -:::warning Enable DEDUP on tables ingested through failover +:::warning Use DEDUP for exactly-once outcomes with failover On unplanned failover — when the primary dies before issuing a durable ACK — the client replays unacknowledged frames against the new primary. Without [DEDUP](/docs/concepts/deduplication/) on the target table, those -replays can produce duplicate rows. Tables ingested through a multi-host -failover connect string **must** declare `DEDUP UPSERT KEYS(...)` covering -row identity. See [Delivery semantics](/docs/concepts/delivery-semantics/) -for the full at-least-once / exactly-once model. +replays can produce duplicate rows. If your application requires exactly-once +outcomes, declare `DEDUP UPSERT KEYS(...)` covering row identity. Applications +that tolerate occasional duplicates can skip DEDUP. See +[Delivery semantics](/docs/concepts/delivery-semantics/) for the full model. ::: @@ -520,9 +520,9 @@ equivalent — same architecture, no durability across restarts. - Taken verbatim. Absolute paths recommended for production; relative paths resolve against the process working directory. - The client does **not** expand shell-style syntax such as `~`. - - Create `sf_dir` before opening the sender unless you use Node.js, which - creates the slot and any missing parent directories recursively. Other - clients may create only the leaf slot directory. + - Java, Rust, C, C++, and Python create `sf_dir` and its slot, but require + any parent directories of `sf_dir` to exist first. Go, .NET, and Node.js + create missing parent directories recursively as well as the slot. - `sender_id` — slot identity. The slot lives at `//`, used verbatim as the directory name. Allowed characters: letters, digits, `_`, `-`. No path separators, no `.`, no spaces. Two senders @@ -672,7 +672,7 @@ exception is a Node.js sender in default memory mode, which gives up after `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies this budget to every outage, and fails with `QwpReconnectExhaustedError` when it runs out. See the - [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). + [Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). - `initial_connect_retry` — whether the client retries the initial connect attempt on failure. - `off` (default, alias `false`) — fail fast on initial connect failure. @@ -754,7 +754,10 @@ transport-level OK ACK alone cannot close. - `durable_ack_keepalive_interval_millis` — interval at which the client emits keepalive PINGs while waiting for durable-ack frames. Required because the server only flushes pending durable acks on inbound recv - events. Default: `200` (ms). Set to `0` or a negative value to disable. + events. Default: `200` (ms). Set to `0` or a negative value to disable + in clients that support it. In Node.js, explicitly setting this key also + requests durable ACK (even at `0`), and negative values are rejected; + see [Node.js differences](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). See the [QWP Egress (WebSocket)](/docs/connect/wire-protocols/qwp-egress-websocket/) wire protocol for the underlying mechanism. diff --git a/documentation/connect/clients/nodejs-operations.md b/documentation/connect/clients/nodejs-operations.md new file mode 100644 index 0000000000..2e5a40b8a6 --- /dev/null +++ b/documentation/connect/clients/nodejs-operations.md @@ -0,0 +1,1413 @@ +--- +slug: /connect/clients/nodejs-operations +title: Node.js client operations and reference +sidebar_label: Node.js operations and reference +description: "Node.js QWP client pool lifecycle, error recovery, failover, configuration, migration, and a complete ingestion and query example." +--- + +For a first connection, row ingestion, and streaming SQL queries, start with the +[Node.js client guide](/docs/connect/clients/nodejs/). This companion page +covers pool lifecycle, concurrency, error recovery, failover, connect-string +differences, migration, and a complete ingestion and query example. + +## The connection pool + +The pooled client keeps two elastic pools: one of senders and one of query +connections. Each pool opens its minimum on `connect()`, grows on demand up to +its maximum, and a housekeeper closes connections that stay idle too long or +exceed their maximum lifetime, never going below the minimum. + +### Borrowing a sender + +A borrowed sender belongs to the borrower until its `close()` returns it: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + for (const [symbol, price] of [ + ["ETH-USD", 2615.54], + ["BTC-USD", 39269.98], + ] as const) { + await sender + .table("trades") + .symbol("symbol", symbol) + .symbol("side", "buy") + .doubleColumn("price", price) + .doubleColumn("amount", 0.1) + .at(Date.now(), "ms"); + } + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence); + } finally { + // Returns the sender to the pool; any pending rows are flushed. + await sender.close(); + } +} finally { + await db.close(); +} +``` + +A long-running producer can keep its borrow for its whole lifetime and call +`flush()` between batches. Size `sender_pool_max` to the number of producers +that hold a sender at the same time. + +`close()` on a borrowed sender flushes its completed rows and returns it to +the pool. The example above waits for the acknowledgement *before* returning +the sender: `close()` itself does not wait for acknowledgements, although in +default memory mode its flush waits for the reconnect during an outage. See +[Closing a borrowed sender](/docs/connect/clients/nodejs/#closing-a-borrowed-sender) for how long that can +take, how to wait for acknowledgements, and what happens when a close fails. + +After `close()`, every method call or property read on that sender object +throws `QwpClientClosedError`. Don't keep references to a returned sender, for +example in callbacks that can run later. + +### Borrowing a query lease + +A query lease runs one query at a time. For concurrent queries, borrow one lease +per query, up to `query_pool_max`: + +```typescript +import { + connectQwpNodeClient, + type QwpQueryLease, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); + +async function countBySymbol(lease: QwpQueryLease, symbol: string) { + const query = await lease.query( + "SELECT count() FROM trades WHERE symbol = $1", + { binds: (binds) => binds.setVarchar(0, symbol) }, + ); + let count = 0n; + for await (const batch of query) count = batch.get(0, 0) as bigint; + await query.completion; + return count; +} + +try { + const [a, b] = await Promise.all([db.borrowQuery(), db.borrowQuery()]); + try { + // Two leases, two WebSockets: the queries run concurrently. + const [eth, btc] = await Promise.all([ + countBySymbol(a, "ETH-USD"), + countBySymbol(b, "BTC-USD"), + ]); + console.log({ eth, btc }); + } finally { + await Promise.all([a.close(), b.close()]); + } +} finally { + await db.close(); +} +``` + +Starting a second query on a lease while one is still active rejects with +`a QWP query is already active on this connection`. Always close a lease in +`finally`: an unreturned lease holds its connection until `db.close()`. + +### Pool settings + +| Key | Default | Purpose | +|---|---|---| +| `sender_pool_min` | `1` | Senders kept open even when idle. `0` lets the pool close them all. | +| `sender_pool_max` | `4` | Maximum senders the pool opens. | +| `query_pool_min` | `1` | Query connections kept open even when idle. | +| `query_pool_max` | `4` | Maximum query connections, which also caps concurrent queries. | +| `acquire_timeout_ms` | `5000` | How long a borrow waits when the pool is at its maximum, before rejecting with `QwpPoolAcquireTimeoutError`. | +| `idle_timeout_ms` | `60000` | Idle time before an excess connection is closed. `0` keeps idle connections. | +| `max_lifetime_ms` | `1800000` | Age at which an idle connection above the pool minimum is closed. Connections kept open by `sender_pool_min` and `query_pool_min` are never recycled, so this does not rotate a pool that is at its minimum. `0` disables it. | +| `housekeeper_interval_ms` | `5000` | How often the housekeeper checks for idle and over-age connections. Minimum `100`. | +| `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | +| `lazy_connect` | `off` | Start without connecting. See below. | + +Pool sizes, acquisition and idle timeouts, lifetime, and housekeeping settings +have typed equivalents in the `pool` section of the second argument +(`senderPoolMin`, `acquireTimeoutMs`, `housekeepingIntervalMs`, and so on). +The other two settings use different locations: + +- `query_close_timeout_ms` maps to `egressSession.cancelDrainTimeoutMs`, not + `pool`. +- Set `lazy_connect=on` in the connect string. When passing a full + `QwpNodeClientOptions` object instead of a string, use top-level + `lazyConnect: true`. It is not supported in `pool` or the second argument. + +When creating a new pooled connection fails, the borrow rejects with +`QwpPoolResourceError`, whose `cause` holds the connection error. + +`borrowSender()` and `borrowQuery()` take no timeout argument. When the pool +is at its maximum, a borrow waits up to `acquire_timeout_ms` for a connection +to be returned. Opening a new connection is bounded by the +[connection timeouts](#connection-timeouts) of each endpoint, and by the +[failover budget](#query-failover) when query retries are on. To enforce a +shorter deadline, such as a request deadline, race the borrow against a timer +and return a lease that arrives late: + +```typescript +import { connectQwpNodeClient, type QwpClient } from "@questdb/nodejs-client"; + +function borrowQueryWithin(db: QwpClient, timeoutMs: number) { + const borrow = db.borrowQuery(); + let timer: ReturnType | undefined; + const deadline = new Promise((_, reject) => { + timer = setTimeout(() => reject(new Error("borrow timed out")), timeoutMs); + }); + return Promise.race([borrow, deadline]) + .catch((error: unknown) => { + // Return a lease that arrives after the deadline. + borrow.then((lease) => lease.close(), () => undefined); + throw error; + }) + .finally(() => clearTimeout(timer)); +} + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await borrowQueryWithin(db, 2_000); + try { + const query = await lease.query("SELECT count() FROM trades"); + for await (const batch of query) console.log(batch.get(0, 0)); + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +### Starting while QuestDB is down + +`connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set +`lazy_connect=on` to start regardless: senders connect in the background and +buffer rows in memory until QuestDB is reachable. The query pool stays empty +until the first query. + +```typescript +import { + connectQwpNodeClient, + QwpIngressAckTimeoutError, +} from "@questdb/nodejs-client"; + +// Resolves immediately, even if QuestDB is not running yet. +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;lazy_connect=on;", +); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); + const sequence = sender.publishedSequence; + // Keep the process running until QuestDB comes back and acknowledges it. + for (;;) { + try { + await sender.waitForAcknowledged(sequence, 10_000); + break; + } catch (error) { + if (!(error instanceof QwpIngressAckTimeoutError)) throw error; + console.info("still waiting for QuestDB; do not restage the row"); + } + } + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +Rows buffered while QuestDB is down exist only in memory. This example keeps +the client running until the row is acknowledged; if the process exits first, +the unacknowledged row may be lost. See +[Closing the pooled client](#closing-the-pooled-client). +To keep them across a shutdown or restart, add a +[store-and-forward](/docs/connect/clients/nodejs/#store-and-forward) journal with `sf_dir`. `lazy_connect=on` +still starts the sender in the background when `sf_dir` is set; the journal +changes where rows are buffered, not whether startup waits for a connection. +Replay from the journal is at least once, so write to a deduplicated table as +described there. + +`lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, +and rejects an explicit conflicting value. Setting `initial_connect_retry=async` +without `lazy_connect` is not enough: the query pool still connects at startup, +so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query +borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. + +A lazy start does not cover two cases. A locked store-and-forward journal +still fails startup; see [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery). And until a +sender has connected once, it cannot check batches against the server's size +limit; see [Batch size limits](/docs/connect/clients/nodejs/#batch-size-limits). + +### Closing the pooled client + +`db.close()` rejects new borrows, then: + +- Cancels active queries and closes every query connection, including leased + ones. +- Closes idle senders. Each publishes its remaining rows and waits up to + `close_flush_timeout_millis` (5 seconds) for QuestDB to acknowledge them. +- Waits for borrowed senders to be returned, until 5 seconds after + `db.close()` was called, or `acquire_timeout_ms` if that is lower. Closing + the idle senders counts toward the same deadline. A sender still borrowed + after that stays open: its owner must `close()` it, and the process stays + alive until then. + +`db.close()` resolves even when an acknowledgement does not arrive in time. +Without `sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they +stay in the journal, and the next sender on the same directory replays them. +The client usually reports the timeout to `ingressSession.onError` as a +non-terminal `QwpIngressAckTimeoutError`, logged as a warning by default, but +the report is best-effort: do not rely on it to detect unacknowledged rows. To +know that QuestDB accepted every row before shutting down, wait for the +acknowledgement before returning each sender (see +[Awaiting acknowledgements](/docs/connect/clients/nodejs/#awaiting-acknowledgements)), or use +[store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). + +## Concurrency + +Node.js runs your code on one thread, but async functions interleave at every +`await`: + +- **`QwpClient`** is safe to share across your whole application. +- **Senders** are not safe for concurrent producers. A row is built across + several calls, so an `await` between `table()` and `at()` lets another task + add columns to the same row. Give each producer its own sender, borrowed from + the pool, and size `sender_pool_max` to match. +- **Query leases** run one query at a time. Borrow one lease per concurrent + query; `query_pool_max` caps concurrent queries. +- **Worker threads** cannot share clients. Create one client per worker, and + give each worker its own `sender_id` when using store-and-forward. + +Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy +work out of them. Row encoding runs on the event loop too, so one process +ingests at most as fast as one CPU core allows. To go faster, split the stream +across worker threads or processes, each with its own client. + +## Error handling + +Each error leaves the client in a known state. The sections after this table +have the details and examples: + +| Error | Surfaces from | State afterwards | What to do | +|---|---|---|---| +| `TypeError`, `RangeError`, or `Error` from local validation | The column method or `at()` that staged the value | The row in progress is discarded; the sender stays usable | Fix the value and write the row again | +| `QwpBatchTooLargeError` | `flush()`, the `at()` whose auto-flush sends the batch, or `close()` | The batch can never be sent: the staged rows are kept, every later flush fails the same way, and `close()` discards them | Call `reset()`, then write the rows again without the oversized one; see [Batch size limits](/docs/connect/clients/nodejs/#batch-size-limits) | +| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, an auto-flushing `at()`, or `close()` | The batch stays staged; the sender stays usable | Keep the sender and flush again later. Don't write the rows again, and don't close the sender while backpressure persists; see [Backpressure](/docs/connect/clients/nodejs/#backpressure) | +| Retriable server rejection | `onSenderError` | The client resends the batch; repeated rejections become terminal | Monitor; no action needed per rejection | +| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection) | +| `QwpReconnectExhaustedError` on a sender | `onError` with `terminal: true`, then the next `flush()`, `at()`, or `close()` | The sender has failed; unsent rows are lost | Borrow a new sender; see [Ingestion reconnect](#ingestion-reconnect) | +| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery) | +| `QwpPoolResourceError` with another `cause` | `connectQwpNodeClient()`, `borrowSender()`, or `borrowQuery()` | No connection was opened | Unwrap `cause`; see [Connection-level errors](#connection-level-errors) | +| `QwpEgressQueryError` | Query iteration and `completion` | The lease stays usable | Fix the SQL or the bind values | +| `QwpEgressQueryTimeoutError`, `QwpEgressQueryAbandonedError` | Query iteration and `completion` | The lease is busy until QuestDB confirms the cancellation | Close the lease and borrow a new one | +| `QwpEgressQueryCancelTimeoutError` | Query iteration and `completion` | The connection is closed | Close the lease and borrow a new one | +| `QwpReconnectExhaustedError` on a query | Query iteration and `completion` | The lease stays failed, even after QuestDB recovers | Close the lease and borrow a new one; see [Query failover](#query-failover) | + +`QwpReplayStoreAppendTimeoutError` and the other store-and-forward journal +errors extend `QwpReplayStoreError`, so test for the specific classes before +the base class, or branch on `error.retryable`. `true` means the failure is +temporary and the sender stays usable. `false` means the journal itself can no +longer be used, for example `QwpReplayStoreCorruptionError` or +`QwpReplayStoreLockLostError`, and the sender has failed. The other classes in +the table have no common base class: test each with `instanceof`. + +### Ingestion errors + +Ingestion reports errors in two ways: + +- **While building a row.** A column method throws, or the promise returned by + `at()` or `atNow()` rejects, with a `TypeError`, `RangeError`, or `Error` for + an invalid value or name. The row in progress is discarded, and the sender + stays usable. +- **Asynchronously, when QuestDB rejects a batch.** The rejection arrives after + `flush()` resolved. It is delivered to the `onSenderError` callback, and + surfaces as a rejection of `waitForAcknowledged()`, or of `flush()` with + `awaitServerAck`. + +```typescript +import { + connectQwpNodeClient, + QWP_SENDER_ERROR_POLICY, + type QwpSenderError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + ingressSession: { + onSenderError: (error: QwpSenderError) => { + // serverStatusByte is absent for client-side errors. + const status = + error.serverStatusByte === undefined + ? "none" + : `0x${error.serverStatusByte.toString(16)}`; + console.error( + `rejected [${error.category}, policy=${error.appliedPolicy}, ` + + `status=${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + + error.serverMessage, + ); + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + // The sender stopped: alert, and fix the data or the schema. + } + }, + }, +}); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", 2615.54) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); + } finally { + // Rethrows a terminal error. The pool then replaces the sender. + await sender.close(); + } +} finally { + await db.close(); +} +``` + +When `onSenderError` is not set, rejections are logged: retriable ones at +`warn`, terminal ones at `error`. Callbacks run asynchronously, never inside the +client's protocol handling, and an exception thrown by a callback is contained. +A standalone sender's `close()` can also reject, with +`QwpSenderCloseTimeoutError`, when its rows are not acknowledged in time; see +[Closing a sender](/docs/connect/clients/nodejs/#closing-a-sender). + +`QwpSenderError` fields: + +| Field | Type | Meaning | +|---|---|---| +| `category` | `string` | `schema-mismatch`, `parse-error`, `security-error`, `write-error`, `internal-error`, `not-writable`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `protocol-violation`, `data-loss`, or `unknown`. Branch on this field. | +| `appliedPolicy` | `string` | What the client did: `retriable` (reconnect and resend), `retriable-other` (resend to another endpoint), `terminal` (the sender stopped), or `abandoned` (journaled data was quarantined). | +| `serverStatusByte` | `number` | The raw QWP status code, for example `0x03` for a schema mismatch. Absent for client-side errors. | +| `serverMessage` | `string` | QuestDB's error text, for example `cannot parse DOUBLE from string [value=abc, column=price]`. | +| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `publishedSequence`. | +| `messageSequence` | `bigint` | The wire sequence of the rejected message. | +| `tableName` | `string` | The table, when the server attributes the rejection to one. Often absent. | +| `detectedAtMs` | `number` | When the client received the rejection. | +| `quarantinedPath` | `string` | For `data-loss` in store-and-forward: where the unreplayable journal was preserved. | + +The default policy follows the category: + +| Category | Policy | Examples | +|---|---|---| +| `schema-mismatch`, `parse-error`, `security-error`, `protocol-violation` | Terminal | Wrong value type for an existing column, malformed data, missing permission | +| `write-error`, `internal-error`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `unknown` | Retriable | Disk pressure, a suspended table, a transient server fault | +| `not-writable` | Retriable on another endpoint | The server is a replica or cannot accept writes | +| `data-loss` | Abandoned | A corrupt store-and-forward journal was set aside | + +A retriable rejection is resent. For rejections that count toward the +poison-frame detector, if the same batch keeps being rejected after +`max_frame_rejections` (4) attempts spanning at least +`poison_min_escalation_window_millis` (5 minutes), the sender stops as for a +terminal error. The `dictionary-gap`, `unknown`, and `not-writable` categories +are exempt: they reset the poison episode instead of adding a strike. +Retriable rejections of symbol-dictionary catch-up frames are also exempt. +The six `on_*_error` connect-string keys are accepted but not applied by this +client. + +Handling notes: + +- **Message stability.** `serverMessage` is free-form English text from the + server. Its wording can change between releases: branch on `category`, not on + the text. +- **Sensitive data.** Server messages can contain column names and values. + Treat them as untrusted input, and redact them before sending them to + third-party error trackers or showing them to end users. +- **Correlation.** There is no server-side request ID. Correlate with the frame + sequence range, `tableName`, and `detectedAtMs`. + +#### Recovering from a terminal rejection + +After a terminal server rejection, the sender is permanently failed. An +already-pending `waitForAcknowledged()` for the rejected batch can reject with +`QwpIngressNackError`. Once the terminal failure is latched, new calls to +`waitForAcknowledged()`, `flush()`, or `close()` reject with +`QwpReplayRejectedError`, whose `status` and message repeat the server's. +Error handlers must allow either class depending on timing. Writing LONG arrays +with `longArrayColumn()` triggers a terminal rejection on every current server. + +Close the sender and create a new one. A pooled sender is replaced +automatically after the `close()` that reports the error. What happens to the +rejected batch depends on the mode: + +- **Without store-and-forward**, the failed sender's unacknowledged batches, + including the rejected one, are discarded with it, and so are rows that a + later borrower staged on it before the error surfaced. The new sender starts + empty. +- **With store-and-forward**, the rejected batch stays at the head of the + journal. Every new sender on that directory, including the pool's + replacement sender and the same client after a restart, sends it again and + fails the same way, with `QwpReplayRejectedError`. Pooled borrows keep + getting that journal, so ingestion through the client stops for every table, + not only the table in the rejected batch, until you act. Treat it as an + outage and alert on it from `onSenderError`. Fix the cause so that QuestDB + accepts the batch, for example by adjusting the table schema, or stop the + process and move the journal directory aside. Moving it aside discards every + unacknowledged batch in it, not only the rejected one. + +### Query errors + +Query errors reject both the `for await` iteration and `completion`: + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, + QwpEgressQueryTimeoutError, +} from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM no_such_table"); + for await (const batch of query) console.log(batch.rowCount); + await query.completion; + } catch (error) { + if (error instanceof QwpEgressQueryError) { + // Prints: 5 [14] table does not exist [table=no_such_table] + console.error(error.status, error.message); + } else if (error instanceof QwpEgressQueryTimeoutError) { + console.error("timed out"); + } else { + throw error; + } + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`QwpEgressQueryError` has `status` (the QWP status code), `message` (the server +text, where parse errors start with the character position in brackets), and +`requestId` (a client-assigned `bigint` that numbers the queries of a +connection). The lease remains usable after a `QwpEgressQueryError`. + +| Status | Name | Meaning | +|---|---|---| +| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column, or a bind value the statement cannot use, such as a boolean for `LIMIT` | +| `0x06` | INTERNAL_ERROR | Execution failure, including a bind value that cannot be converted, such as `'abc'` compared with a DOUBLE column, and a column type that QWP cannot return | +| `0x08` | SECURITY_ERROR | Missing permission | +| `0x0a` | CANCELLED | The query was cancelled with `cancel()` | +| `0x0b` | LIMIT_EXCEEDED | A server limit was reached: the server-side query timeout, memory, or a result row too large to send | + +The `QWP_STATUS` export names these codes, for example +`QWP_STATUS.PARSE_ERROR`, so code can compare against constants instead of +numbers. A status alone does not separate a client mistake from a server +fault: `0x06` covers both bind values that cannot be converted and execution +failures, and `0x0b` covers both the server-side query timeout and memory +limits. + +Other query errors: + +| Error | Meaning | +|---|---| +| `QwpEgressQueryTimeoutError` | The query deadline expired and cancellation started. Has `requestId` and `timeoutMs`. | +| `QwpEgressQueryAbandonedError` | Iteration ended early, for example with `break`. | +| `QwpEgressQueryCancelTimeoutError` | QuestDB did not confirm a cancellation in time; the connection was closed. | +| `QwpEgressSessionClosedError` | The query connection is closed. | +| `QwpReconnectExhaustedError` | Failover gave up; see [Query failover](#query-failover). | + +As with ingestion, the message text is not stable, may echo parts of the SQL, +and has no server-side correlation ID beyond `requestId`. + +### Connection-level errors + +| Error | Raised when | +|---|---| +| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, or `transport`. It also carries `statusCode`, `retryable`, and `url`. | +| `QwpFailoverError` | Every endpoint in a multi-host list failed. `attempts` holds each endpoint and its error. | +| `QwpPoolResourceError` | The pool could not open a new connection. `cause` holds the error above. | +| `QwpPoolAcquireTimeoutError` | Every pooled connection stayed leased past `acquire_timeout_ms`. | +| `QwpReconnectExhaustedError` | The reconnect budget ran out. The sender or query failed permanently. | +| `QwpRoleMismatchError` | No endpoint has the role that `target` requires. | +| `QwpDurableAckUnavailableError` | `request_durable_ack=on`, but the server does not support it. | +| `QwpClientClosedError` | The pooled client, or a returned lease, is already closed. | + +The pooled client wraps every failure to open a connection, from +`connectQwpNodeClient()`, `db.connect()`, `borrowSender()`, or +`borrowQuery()`, in a `QwpPoolResourceError`. Unwrap its `cause` before +checking for a specific error. When `addr` lists several hosts, the cause is a +`QwpFailoverError` whose `attempts` hold the error of each endpoint. When +initial-connect retry is on, for example with a `failover_*` or `reconnect_*` +key, or with a typed `egressSession.reconnect` object for query connections +(see [Typed reconnect policy](#typed-reconnect-policy)), the cause is a +`QwpReconnectExhaustedError` instead, and its own `cause` holds the last +attempt's error: + +```typescript +import { + connectQwpNodeClient, + QwpFailoverError, + QwpPoolResourceError, + QwpReconnectExhaustedError, + QwpUpgradeError, +} from "@questdb/nodejs-client"; + +// The errors behind a failed connection, one per endpoint tried. +function connectionErrors(error: unknown): unknown[] { + let cause = error instanceof QwpPoolResourceError ? error.cause : error; + // With initial-connect retry on, the last attempt's error is wrapped. + if (cause instanceof QwpReconnectExhaustedError) cause = cause.cause; + return cause instanceof QwpFailoverError + ? cause.attempts.map((attempt) => attempt.error) + : [cause]; +} + +try { + const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); + await db.close(); +} catch (error) { + for (const cause of connectionErrors(error)) { + if (cause instanceof QwpUpgradeError && cause.kind === "authentication") { + console.error("QuestDB rejected the credentials:", cause.message); + } else { + console.error("cannot connect:", cause); + } + } + throw error; +} +``` + +An authentication rejection (HTTP 401 or 403) is terminal before a sender's +first successful connection and for query connections. It stops the endpoint +walk because credentials are assumed to be shared across the cluster. + +After a successful connection, regular senders with `sf_dir` or in background +memory mode (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying +authentication rejections indefinitely. This lets buffered data drain once +server-side authentication is restored. Senders in default memory mode, and +orphan drainers, do not have this exception. See +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) +for how other clients behave. + +Endpoints in error messages have any embedded credentials removed. + +#### Connection timeouts + +Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS +and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and +authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits +`connect_timeout` when only the latter is set. A timeout in either phase +produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or +`authentication`. + +After the upgrade, a query connection has a separate 5-second deadline for +the initial QWP `SERVER_INFO` frame. Configure it with the typed option +`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not +change it. Expiry produces an ordinary `Error` with the message +`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. + +### Logging + +The client writes its own messages to the console by default, at the `error`, +`warn`, and `info` levels. To route a sender's messages, such as warnings about +rows discarded on close, through your logger, pass a `QwpSenderLogger` +function: `{ sender: { log } }` as the second argument of +`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. Its signature +is `(level: "error" | "warn" | "info" | "debug", message: string | Error)`, so +convert `message` with `String()` if your logger takes strings only. The +function also receives `debug` messages, one per staged row, so filter by +level. + +Rejected batches and session errors go to `ingressSession.onSenderError` and +`ingressSession.onError`. Their defaults log to the console, so replace both to +route them through your logger. Some messages from other parts of the client, +such as store-and-forward recovery, always go to the console. + +## Failover and high availability + +:::note Enterprise + +Failing over between several QuestDB hosts requires QuestDB Enterprise +replication. Reconnecting to a single restarted server works in open source +too. + +::: + +### Multiple endpoints + +List several hosts in `addr`: + +```text +wss::addr=db-a.example.com:9000,db-b.example.com:9000,db-c.example.com:9000; +``` + +The client ranks endpoints by observed health and by `zone`, and on a +connection loss moves to the next usable one. `addr` is shared by ingestion and +queries. + +Ingestion always needs the primary: replicas refuse writes, and the sender +walks the list until it finds the current primary. Queries can use any node. +`target` selects which roles queries accept: `any` (the default), `primary`, or +`replica`. Set it with the typed `egress` option, as below, because in the +connect string `target` also filters ingestion (see the caution that follows). + +`target` is a strict filter, not a preference. With `replica`, queries never +fall back to the primary, and they fail when no replica is reachable, including +against a single open source server. Because the pooled client opens a query +connection at startup, `connectQwpNodeClient()` then fails too, with a +`QwpPoolResourceError` whose `cause` leads to a `QwpRoleMismatchError`; see +[Connection-level errors](#connection-level-errors) to unwrap it. To start +without a replica, also set `query_pool_min=0`. Queries borrowed before a +replica is reachable then reject with `QwpPoolResourceError`: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// Queries run on replicas only. Ingestion still follows the primary. +const db = await connectQwpNodeClient( + "wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN;" + + // Start, and ingest, even while no replica is reachable. + "query_pool_min=0;", + { egress: { target: "replica" } }, +); +await db.close(); +``` + +`zone` prefers endpoints in the same zone. + +:::caution `target` in the connect string also filters ingestion + +Unlike the Java client, the Node.js client applies `target` and `zone` from +the connect string to ingestion as well as queries. `target=replica` in the +connect string therefore stops ingestion from reaching the primary. To read +from replicas and write to the primary with one client, keep `target` out of +the connect string and set it for queries only: +`connectQwpNodeClient(conf, { egress: { target: "replica" } })`. + +::: + +### Ingestion reconnect + +When the connection drops, the sender reconnects with exponential backoff and +jitter, then resends every unacknowledged batch: + +| Key | Default | Purpose | +|---|---|---| +| `reconnect_initial_backoff_millis` | `100` | First retry delay. | +| `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | +| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in default memory mode. `0` removes the limit. | +| `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | + +Whether the sender gives up depends on the [ingestion mode](/docs/connect/clients/nodejs/#ingestion-modes): + +- **Default memory mode** retries for up to `reconnect_max_duration_millis` + per outage. When the budget runs out, the sender fails permanently with + `QwpReconnectExhaustedError`, and its unsent rows are lost. The Java + reference client retries indefinitely in this mode instead. +- **Background memory mode** (`initial_connect_retry=async` or + `lazy_connect=on`) and **store-and-forward** (`sf_dir`) retry indefinitely. + +Setting any `reconnect_*` key also makes a sender's first connection retry +within the budget, as if `initial_connect_retry=on`. Set +`initial_connect_retry=off` explicitly to keep a fail-fast start. The keys do +not apply to query connections: the pooled client still opens its query pool +at startup, so `connectQwpNodeClient()` fails fast while QuestDB is down unless +you also set `query_pool_min=0` or enable query retries (see +[Connection events](#connection-events)). + +Replay after a reconnect is at least once: a batch that QuestDB committed just +before the connection dropped is sent again. Write to a deduplicated table, as +described under [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward), to keep replayed rows +from inserting duplicates. + +### Query failover + +If the connection fails during a query, the client reconnects, to another +endpoint when there is one, and runs the query again from the start: + +| Key | Default | Purpose | +|---|---|---| +| `failover` | `on` | Set `off` to fail the query instead of retrying. | +| `failover_max_attempts` | `8` | Connection attempts per failure. Each attempt tries every endpoint in `addr`. | +| `failover_backoff_initial_ms` | `50` | First retry delay. | +| `failover_backoff_max_ms` | `1000` | Longest delay between retries. | +| `failover_max_duration_ms` | `30000` | Time budget per failure. | + +The attempt limit and the time budget apply together, and whichever is reached +first ends the failover. When attempts fail fast, for example with connection +refused while a server restarts, the 8 attempts and their backoff of 50 ms to +1 second, with jitter, take only about 1 to 3 seconds, and at most about 4.5 +seconds, long before the 30-second budget. To ride out a longer restart, raise `failover_max_attempts`, +or set `maxAttempts: 0` in a typed `egressSession.reconnect` object to remove +the attempt limit and rely on the time budget alone; see +[Typed reconnect policy](#typed-reconnect-policy). + +When failover gives up, the query rejects with `QwpReconnectExhaustedError`, +and the lease stays failed even after QuestDB recovers: close it and borrow a +new one. A `QwpEgressQueryError` from the server is a query result and never +triggers failover. Replaying an in-flight `query()` also re-executes DDL and +DML: an `INSERT` may run twice if its completion was lost. For non-idempotent +SQL, use a separate client configured with `failover=off` and check an +uncertain outcome before retrying; see +[DDL and DML statements](/docs/connect/clients/nodejs/#ddl-and-dml-statements). + +:::warning Clear partial results when a query restarts + +A re-executed query starts again from the first row. Batches that were queued +but not yet consumed are discarded for you, but rows your loop already +processed are delivered again. If your code accumulates rows, clear them when +the query restarts; otherwise it sees the first part of the result twice. + +::: + +Every batch has a `batchSequence` starting at `0n`, including the first batch +after a replay. Clear accumulated rows on that batch. A replay may return +**zero rows and no batches**, though, leaving prior rows in your accumulator. +`egressSession.onReplayReset` also clears it when that happens. Since this +callback is shared across the pool and request IDs are per connection, the +example limits the pool to one active query: + +```typescript +import { + connectQwpNodeClient, + QwpReconnectExhaustedError, +} from "@questdb/nodejs-client"; + +const rows: (readonly unknown[])[] = []; +const db = await connectQwpNodeClient( + "ws::addr=db-a.example.com:9000,db-b.example.com:9000;query_pool_max=1;", + { + egressSession: { + onReplayReset: () => { + rows.length = 0; + }, + }, + }, +); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT * FROM trades LIMIT 100000", { + // The deadline covers the whole query, including a re-execution. + timeoutMs: 30_000, + }); + for await (const batch of query) { + // Sequence 0 starts the result, both initially and after a failover. + if (batch.batchSequence === 0n) rows.length = 0; + for (const row of batch.rows()) rows.push(row); + } + await query.completion; + console.log(`${rows.length} rows`); + } catch (error) { + if (!(error instanceof QwpReconnectExhaustedError)) throw error; + // Failover gave up. This lease stays failed: return it, retry later. + console.error("no endpoint could run the query:", error.message); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +The reset needs the rows of the current attempt in one place you can discard. +When your code passes rows on as they arrive, for example by streaming them to +an HTTP response, it cannot take them back after a restart. Then choose one +of these instead: + +- Run such queries on a separate client with `failover=off`, so that a lost + connection fails the query instead of restarting it, and retry the whole + request. +- Keep each result small enough to buffer, for example by paging with + `WHERE timestamp < $1 ORDER BY timestamp DESC LIMIT n`, binding the oldest + timestamp of the previous page. +- Treat `batchSequence === 0n` after rows have left as an error, and abort the + downstream response instead of sending duplicates. + +`egressSession.onReplayReset` runs before a query is replayed, including when +that replay returns no batches. Its event has `requestId`, `endpoint`, +`previousEndpoint`, `serverInfo`, and `cause`. The `requestId` matches +`query.requestId`, but request IDs are numbered per connection and every lease +of a pooled client shares the callback: it cannot identify which of several +concurrent queries restarted. Use a dedicated, single-query client when the +callback resets result state, as above. For concurrent results that cannot be +isolated, set `failover=off` and retry the whole query after a transport error; +`batchSequence === 0n` alone cannot detect a zero-batch replay. + +### Typed reconnect policy + +Reconnect and failover behavior comes from the connect-string keys above, or +from typed `reconnect` objects in the second argument of +`connectQwpNodeClient()`: `ingressSession.reconnect` for senders and +`egressSession.reconnect` for queries. You need the object to register +`onEvent` for [connection events](#connection-events). + +:::caution A typed `reconnect` object replaces the connect-string keys + +`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces +the whole reconnect policy parsed from the `reconnect_*`, +`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and +`egressSession.reconnect` replaces the policy parsed from `failover*` keys. +Fields you leave out of the object take the built-in defaults, not the values +from the connect string. When you supply the object, for example to register +`onEvent`, set every bound you rely on in it. + +::: + +The object's fields and the connect-string keys they replace: + +| Field | Ingestion key, default | Query key, default | +|---|---|---| +| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8`. The key accepts `1` or more; the typed field also accepts `0`, unlimited | +| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | +| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | +| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | +| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | +| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | +| `onEvent` | None | None | + +`egressSession: { reconnect: false }` is the typed equivalent of +`failover=off`. Senders in background memory mode or with `sf_dir` retry +indefinitely: `maxAttempts` and `maxDurationMs` do not end their retries, so +the full example's `reconnect: { onEvent }` with `sf_dir` keeps retrying +through an outage of any length. + +The two directions treat the first connection differently: + +- **Senders**: setting any `reconnect_*` key makes the first connection retry + within the budget, as if `initial_connect_retry=on`. A typed + `ingressSession.reconnect` object does not, so the first connection still + fails fast. Set `initial_connect_retry` in the connect string to choose the + startup behavior. +- **Query connections**: the first connection retries within the failover + budget, for retryable errors, when you supply an `egressSession.reconnect` + object, set `failover=on` explicitly, or set a `failover_*` key without + `failover=off`. Otherwise it is attempted once. `failover=off` and + `egressSession.reconnect: false` turn reconnects off entirely. + +### Connection events + +Register `reconnect.onEvent` to observe connections. Events are delivered +asynchronously through a bounded queue (64 by default). The connect-string +`connection_listener_inbox_capacity` key configures the ingestion queue only; +for query events, set the typed `egressSession.connectionListenerInboxCapacity` +option. When a queue overflows, its oldest events are dropped and counted in +the metrics. + +```typescript +import { + connectQwpNodeClient, + QWP_RECONNECT_EVENT_KIND, + type QwpReconnectEvent, +} from "@questdb/nodejs-client"; + +function onEvent(event: QwpReconnectEvent) { + switch (event.kind) { + case QWP_RECONNECT_EVENT_KIND.RECONNECTING: + console.warn("connection lost, reconnecting:", event.cause); + break; + case QWP_RECONNECT_EVENT_KIND.FAILED_OVER: + console.warn(`failed over to ${String(event.endpoint)}`); + break; + default: + console.info(event.kind, String(event.endpoint ?? "")); + } +} + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + // Each object replaces the reconnect_* or failover* keys from the connect + // string. Fields left out use the built-in defaults. + ingressSession: { + reconnect: { onEvent }, + // No event marks a terminal failure: it arrives here instead. + onError: (event) => { + if (event.terminal) console.error("ingestion stopped:", event.error); + }, + }, + egressSession: { reconnect: { onEvent } }, +}); +await db.close(); +``` + +Supplying the `reconnect` objects also changes how the first connection is +retried; see [Typed reconnect policy](#typed-reconnect-policy). + +| Kind | Meaning | +|---|---| +| `connected` | The first connection succeeded. | +| `reconnecting` | The active connection was lost. `cause` holds the error. | +| `attempt-failed` | One connection attempt failed. The client retries if the error is retryable and its budget allows; otherwise this is the last event before the failure is reported. | +| `reconnected` | Reconnected to the same endpoint. | +| `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | +| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. A background-started sender emits this from startup, including with `sf_dir`; a foreground-started sender with `sf_dir` retries after its first successful connection. | +| `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | +| `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | + +The `QWP_RECONNECT_EVENT_KIND` constants name these kinds: `CONNECTED`, +`RECONNECTING`, `ATTEMPT_FAILED`, `RECONNECTED`, `FAILED_OVER`, +`DURABLE_ACK_UNAVAILABLE`, `DURABLE_ACK_PERSISTENT_FAILURE`, and +`PRIMARY_UNAVAILABLE`. + +`reconnected` and `failed-over` are mutually exclusive: code that tracks the +current node must handle both. The client has no property that says whether it +is connected right now. To report it, for example in a health check, track the +latest event: after `reconnecting` the connection is down, and `connected`, +`reconnected`, or `failed-over` mean it is up. A query that reconnects runs +again from its first batch; see +[Query failover](#query-failover) for resetting accumulated rows. + +No event marks a terminal failure. When a sender stops retrying, because its +reconnect budget ran out or the error cannot be retried, +`ingressSession.onError` receives a `QwpIngressErrorEvent` with +`terminal: true`, even while the sender is idle. The event also has `error`, +`timestampMs`, and, for a server rejection, `senderError`. The sender's next `flush()`, auto-flushing `at()`, or +`close()` then rejects with the same error, such as +`QwpReconnectExhaustedError`. A query that cannot fail over rejects its +iteration and `completion` instead. + +For ingestion, `ingressSession` also accepts `onProgress`, for published, +acknowledged, and durably acknowledged sequences, and `onError`, for session +errors. `sender.metrics` returns a snapshot of the sender's counters, including +`metrics.ingress` with the replay queue, reconnect, and notification counters. + + + +## Configuration reference + +The [connect string reference](/docs/connect/clients/connect-string/) documents +every key. The Node.js client's defaults and deviations: + +| Key | Default | Notes | +|---|---|---| +| `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | +| `username`, `password`, `token` | none | Basic or bearer authentication. | +| `tls_verify`, `tls_roots` | `on`, Node.js CA bundle | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | +| `connect_timeout`, `auth_timeout_ms` | `15000` | DNS and TCP/TLS connection, and upgrade deadlines, in milliseconds. See [Connection timeouts](#connection-timeouts). | +| `auto_flush` | `on` | Master switch for the three triggers. | +| `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | +| `auto_flush_interval` | `100` | Milliseconds. `0` disables. `off` is rejected. | +| `auto_flush_bytes` | disabled | Size, or `off`. | +| `close_flush_timeout_millis` | `5000` | ACK wait in a standalone sender's `close()`. | +| `transaction` | `off` | Keep auto-flushed batches in an open transaction until `flush()`. | +| `request_durable_ack`, `durable_ack_keepalive_interval_millis` | `off`, `200` | Enterprise. Explicitly setting the keepalive interval alone requests durable ACK; a negative interval is rejected. | +| `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | +| `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | +| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in default memory mode. `0` removes it. | +| `max_frame_rejections`, `poison_min_escalation_window_millis` | `4`, `300000` | Poison-frame detector: rejections of one batch, and the minimum time they must span, before the sender stops. | +| `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | +| `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | +| `sf_durability` | `memory` | `memory`, `periodic`, or `append`. Requires `sf_dir` when explicitly set, even to `memory`. | +| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](/docs/connect/clients/nodejs/#sf-capacity), not a hard disk limit; memory queue cap without `sf_dir`. | +| `sf_max_segment_bytes` | `4m` with `sf_dir`, none without | Journal segment size, which also caps a batch. | +| `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | +| `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. Both require `sf_dir` when explicitly set. | +| `target`, `zone` | `any`, none | Endpoint role and zone preference. Apply to ingestion too. | +| `failover`, `failover_max_attempts`, `failover_max_duration_ms` | `on`, `8`, `30000` | Query failover. | +| `compression`, `compression_level` | `raw`, `1` | Query result compression. Explicit `compression_level` requires `compression=zstd` or `auto`. | +| `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | +| `client_id` | `typescript/` | Sent to the server for diagnostics. | +| `error_inbox_capacity`, `connection_listener_inbox_capacity` | `256`, `64` | Queues for rejection callbacks and ingestion connection events. Query event queue: typed `egressSession.connectionListenerInboxCapacity`. | +| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client. A standalone `Sender` also applies `lazy_connect`. | + +The +[API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +covers every type and option. The +[QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) +in the client repository describes the delivery semantics in depth. + +### Programmatic options + +Callbacks, custom agents, and other settings a string cannot express go in the +second argument, a `QwpNodeClientConfigOptions` object. When the connect string +and typed options set the same option, the typed value wins. Credentials and +TLS are the exception: a typed `webSocket.authorization` header cannot be +combined with `token`, `username`, or `password` in the string, and a typed +`webSocket.agent` cannot be combined with `tls_verify` or `tls_roots`. Both +combinations are rejected. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + // QwpSenderOptions for every pooled sender + sender: { awaitServerAck: true }, + // Ingestion callbacks and replay settings + ingressSession: { + onSenderError: (error) => + console.error("rejected batch", error.category, error.serverMessage), + }, + // Query session defaults + egressSession: { + queryTimeoutMs: 30_000, + cancelDrainTimeoutMs: 5_000, + serverInfoTimeoutMs: 10_000, + }, + // Egress-only routing and compression + egress: { compression: "zstd" }, + // Pool sizes and timeouts + pool: { senderPoolMax: 2, queryPoolMax: 8 }, +}); +await db.close(); +``` + +The typed `egress` section takes `target`, `zone`, `compression`, +`compressionLevel`, and `maxBatchRows`. The other sections are `webSocket` +(connection settings shared by both directions, such as `agent` or +`connectTimeoutMs`) and `storeAndForward`, the journal settings described +under [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). Its fields match the connect +string keys: + +| Typed field | Connect string key | +|---|---| +| `directory` | `sf_dir` | +| `maxBytes` | `sf_max_total_bytes` | +| `maxSegmentBytes` | `sf_max_segment_bytes` | +| `durability` | `sf_durability` | +| `checkpointIntervalMs` | `sf_sync_interval_millis` | +| `appendDeadlineMs` | `sf_append_deadline_millis` | +| `drainOrphans` | `drain_orphans` | +| `maxBackgroundDrainers` | `max_background_drainers` | + +The second argument has no field for `sender_id`; set it in the connect +string. + +`Sender.fromConfig()` takes `{ log, agent, qwp }` as its second argument, +where `qwp` has the sections `webSocket`, `session` (the equivalent of +`ingressSession`), `sender`, and `udp`. + +The `reconnect` objects in `ingressSession` and `egressSession` replace the +whole reconnect policy parsed from the connect string; see +[Typed reconnect policy](#typed-reconnect-policy) before you set one. + +### Differences from other clients + +The Node.js client differs from the Java reference client, and from the shared +[connect string reference](/docs/connect/clients/connect-string/), in these +places: + +| Area | Node.js behavior | +|---|---| +| Outage budget | A sender in default memory mode gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. Senders in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) or with `sf_dir` retry indefinitely. See [Ingestion reconnect](#ingestion-reconnect). | +| `target` and `zone` | Also apply to ingestion. Set a query-only role with the typed `egress.target` option. See [Multiple endpoints](#multiple-endpoints). | +| Authentication rejected after a first connection | Senders with `sf_dir` or in background memory mode keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | +| Durable acknowledgement unavailable | Senders started in the background retry from startup even with `sf_dir`; foreground store-and-forward senders fail on first connect, but retry after a successful connection. They emit `durable-ack-unavailable` while retrying. See [Durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). | +| `sf_durability` and SF-only keys | Also accepts `append`, but rejects explicit `sf_durability` (even `memory`), `sf_sync_interval_millis`, `drain_orphans` (even `off`), `max_background_drainers`, or `catch_up_cap_gap_min_escalation_window_millis` without `sf_dir`. Unlike Java, do not pass these keys in a memory-mode shared string. | +| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Journal capacity](/docs/connect/clients/nodejs/#sf-capacity). | +| `sf_dir` path creation | Creates missing parent directories and the slot recursively; Java and Rust-derived clients only create `sf_dir` and the slot. | +| `durable_ack_keepalive_interval_millis` | Explicitly setting this key, even to `0`, also requests durable ACK; it fails against OSS if the sender connects. Negative values throw `RangeError` (the shared reference treats them as disabled). | +| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery). | +| `max_lifetime_ms` | Closes idle connections above the pool minimum only. Connections at the minimum are not recycled. | +| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush or from sender creation. Size values take single-letter suffixes only. `compression_level` requires `compression=zstd` or `auto`. `tls_roots` must be PEM; `tls_roots_password`, `init_buf_size`, and `max_buf_size` are rejected. | +| Initial connection and reconnect | `lazy_connect=on` opens senders in the background at startup instead of waiting for a borrow. A query's first connection retries with explicit `failover=on`, a `failover_*` key (unless `failover=off`), or typed `egressSession.reconnect`, not just from the default `failover=on`. See [Starting while QuestDB is down](#starting-while-questdb-is-down) and [Typed reconnect policy](#typed-reconnect-policy). Ingestion reconnect uses full jitter (delay from 0 up to the backoff ceiling), not the equal-jitter schedule in the shared failover guide. | +| `connect_timeout` | Also covers DNS and the TLS handshake, and `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. See [Connection timeouts](#connection-timeouts). | +| `tls_roots` default | The CA certificates bundled with Node.js, not the operating system's trust store. See [TLS](/docs/connect/clients/nodejs/#tls). | +| Defaults | `connect_timeout` is `15000` and `poison_min_escalation_window_millis` is `300000`. `close_flush_timeout_millis` is `5000`, as in the Rust, C, C++, Python, and Go clients; Java and .NET use `60000`. | +| Close after an ACK timeout | A standalone sender's `close()` rejects with `QwpSenderCloseTimeoutError` instead of logging a warning. The pooled client's `db.close()` resolves and reports the timeout, best-effort, to `ingressSession.onError`. See [Closing a sender](/docs/connect/clients/nodejs/#closing-a-sender). | +| Error reports | Categories and policies are lowercase, hyphenated strings, such as `schema-mismatch` and `retriable-other`. See [Ingestion errors](#ingestion-errors). | +| Pool and query keys on a standalone `Sender` | The `Sender` logs a warning for the pool and query-only keys it ignores. It applies `client_id` and `lazy_connect`. | +| `connection_listener_inbox_capacity` | Sets the ingestion event inbox only. Set the query inbox with typed `egressSession.connectionListenerInboxCapacity`; see [Connection events](#connection-events). | +| `on_*_error` keys | Accepted but not applied. | + +## Migration + +### From ILP to QWP + +The row API is unchanged, so existing `Sender` code migrates by changing the +connect string and calling `connect()`: + +```diff +- const sender = await Sender.fromConfig("http::addr=localhost:9000"); ++ const sender = await Sender.fromConfig("ws::addr=localhost:9000"); ++ await sender.connect(); +``` + +| Aspect | ILP over HTTP | QWP over WebSocket | +|---|---|---| +| Connect string schema | `http::`, `https::` | `ws::`, `wss::` | +| Auto-flush rows | 75,000 (600 over TCP) | 1,000 | +| Auto-flush interval | 1,000 ms | 100 ms | +| `flush()` completes when | QuestDB responds to the HTTP request | The batch is published; the ACK arrives later | +| Server rejection | `flush()` throws | Asynchronous: `onSenderError`, `waitForAcknowledged()`, or `flush()` with `awaitServerAck` | +| Rows staged at `close()` | Lost unless flushed | Published; waits up to 5 seconds for ACK, then unacknowledged rows may be lost without `sf_dir` | +| Reconnect and replay | Retries one request for `retry_timeout` | Automatic, with replay of unacknowledged batches | +| Store-and-forward, querying, pooling | Not available | Available | +| Column types | ILP types | More types, subject to [column-method](/docs/connect/clients/nodejs/#column-methods) and [array](/docs/connect/clients/nodejs/#arrays) support | + +Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, +`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss` +with a hint. It names the replacement where there is one +(`retry_timeout` becomes `reconnect_max_duration_millis`, and `tls_ca` becomes +`tls_roots`), and otherwise says that the key applies only to ILP or that QWP +negotiates the setting itself. To keep ILP-sized batches, set +`auto_flush_rows` and `auto_flush_interval` explicitly. Migrate one sender at a +time: ILP and QWP senders can run side by side. + +### Upgrading from 4.x + +Version 5.0.0 keeps the ILP API and adds QWP. Changes that affect existing ILP +code: + +- **Null values.** Passing `null` or `undefined` to a column or symbol method now + omits the column. Existing nullable columns store NULL; BOOLEAN defaults to + `false`, and BYTE and SHORT default to `0` (see [Null values](/docs/connect/clients/nodejs/#null-values)). + Earlier versions threw a type error for most such values. Validate data + before calling the sender if you relied on the error. +- **Decimal scale.** `decimalColumn()` over ILP rejects a non-integer `scale` + with a `RangeError`. Earlier versions silently coerced it, writing `2.5` as + scale 2 and `NaN` as scale 0. +- **`intColumn()`** also accepts a `bigint`, for LONG values beyond + `Number.MAX_SAFE_INTEGER`. +- **TCP authentication** now works on Node.js 26, which rejects the JWK the + client previously built. +- **New dependency.** The package now depends on `ws`, used for QWP. + +## ILP transports (legacy) + +The Node.js `Sender` still ingests over ILP, for existing deployments and for +servers without QWP. To move ILP code to QWP, see +[From ILP to QWP](#from-ilp-to-qwp); for behavior changes in 5.0.0, see +[Upgrading from 4.x](#upgrading-from-4x). ILP senders support HTTP (`http::`, +`https::`) and TCP (`tcp::`, `tcps::`) transports: + +```typescript +import { Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig( + "http::addr=localhost:9000;username=admin;password=quest;", +); +try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.00044) + .at(Date.now(), "ms"); + // ILP does not flush on close: rows still buffered at close() are lost. + await sender.flush(); +} finally { + await sender.close(); +} +``` + +- HTTP connects per request, so `connect()` is not needed; TCP transports + require `await sender.connect()`. `token=...` selects bearer authentication + over HTTP. Over TCP, `username` and `token` set the JWK key ID and private + key. +- Over HTTP, `flush()` sends the buffer as one request and throws if QuestDB + rejects it. Data is transactional only for a single-table request. A + multi-table request can commit earlier tables before a later table fails, + so a failed flush does not mean no data was committed. Schema changes, such + as automatically added columns, are not rolled back even for a single-table + request. See [HTTP transaction semantics](/docs/connect/compatibility/ilp/overview/#http-transaction-semantics). +- Decimals need ILP protocol version 3: HTTP negotiates it automatically, and + TCP needs `protocol_version=3`. Arrays need version 2 or later. +- Undici is the default HTTP agent. Set `stdlib_http=on` to use the Node.js + `http` module instead. + +For ILP options, see the +[`SenderOptions` reference](https://questdb.github.io/nodejs-questdb-client/classes/_questdb_nodejs-client.SenderOptions.html) +and the [ILP overview](/docs/connect/compatibility/ilp/overview/). + +## Full example: Ingestion and querying with failover + +A production-oriented pattern that ingests trades and queries recent prices, +with TLS, a token, several hosts, error handling, and failover handling. Before +running it, create the deduplicated table on the primary (or reuse the table +from [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward)). If the table is missing, QWP +creates it without deduplication: + +```questdb-sql +CREATE TABLE IF NOT EXISTS trades_sf ( + timestamp TIMESTAMP, + trade_id VARCHAR, + symbol SYMBOL, + side SYMBOL, + price DOUBLE, + amount DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, trade_id); +``` + +Replace the sample events with source-assigned trade IDs and timestamps. Keep +both values unchanged when retrying the same event, and use a writable, +persistent `sf_dir` so unacknowledged rows survive a shutdown: + +```typescript +import { + connectQwpNodeClient, + QwpEgressQueryError, + QwpIngressAckTimeoutError, + QwpPoolResourceError, + QWP_RECONNECT_EVENT_KIND, + QWP_SENDER_ERROR_POLICY, + type QwpReconnectEvent, + type QwpSenderError, +} from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); + +// This example runs one query at a time so its replay callback can clear it. +const recentPrices: (readonly unknown[])[] = []; + +// Replace with your alerting. +function alertOperator(message: string) { + console.error("ALERT:", message); +} + +function logConnection(event: QwpReconnectEvent) { + if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { + const endpoint = String(event.endpoint ?? ""); + console.info("questdb connection:", event.kind, endpoint); + } +} + +const db = await connectQwpNodeClient( + "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + + `token=${token};` + + // append: every flush waits for a disk sync; see "Store-and-forward". + "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + + // Limit offline batches below the default server's 2 MiB limit. + "sf_durability=append;sf_max_segment_bytes=1m;sender_pool_max=4;" + + // Query pool stays cold while replicas are down; one active query so the + // replay callback below can reset its state even if replay has no batches. + "query_pool_min=0;query_pool_max=1;", + { + // Queries run on replicas only, never on the primary; ingestion always + // follows the primary. + egress: { target: "replica", compression: "zstd" }, + ingressSession: { + onSenderError: (error: QwpSenderError) => { + console.error("batch rejected:", error.category, error.serverMessage); + // A terminally rejected batch stays in the journal and blocks + // ingestion through this client, for every table, until it is fixed. + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + alertOperator(`QuestDB rejected a batch: ${error.serverMessage}`); + } + }, + // Terminal failures, such as a batch that QuestDB rejects terminally. + onError: (event) => { + if (event.terminal) console.error("ingestion stopped:", event.error); + }, + // Replaces any reconnect_* keys; omitted fields use the defaults. + reconnect: { onEvent: logConnection }, + }, + egressSession: { + queryTimeoutMs: 30_000, + // Replaces any failover* keys; omitted fields use the defaults. + // No 8-attempt limit (maxAttempts 0): failover can last up to 30 s. + reconnect: { + maxAttempts: 0, + maxDurationMs: 30_000, + onEvent: logConnection, + }, + onReplayReset: (event) => { + recentPrices.length = 0; + console.warn("query restarts on", String(event.endpoint)); + }, + }, + }, +); + +try { + // Ingestion: one borrowed sender per producer. IDs and timestamps must + // come from the source, not be regenerated on an application retry. + const events = [ + { + tradeId: "trade-12345", + timestampMs: 1723000000000, + symbol: "ETH-USD", + price: 2615.54, + amount: 0.5, + }, + { + tradeId: "trade-12346", + timestampMs: 1723000000001, + symbol: "BTC-USD", + price: 39269.98, + amount: 0.001, + }, + ]; + const sender = await db.borrowSender(); + try { + for (const event of events) { + await sender + .table("trades_sf") + .stringColumn("trade_id", event.tradeId) + .symbol("symbol", event.symbol) + .symbol("side", "buy") + .doubleColumn("price", event.price) + .doubleColumn("amount", event.amount) + .at(event.timestampMs, "ms"); + } + await sender.flush(); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); + } catch (error) { + if (!(error instanceof QwpIngressAckTimeoutError)) throw error; + console.warn("ACK timed out; rows remain in sf_dir for replay after close"); + } finally { + // After a terminal rejection, close() rejects with the same failure. + // Log it so that it does not replace the error thrown above. + await sender + .close() + .catch((error) => console.error("close failed:", error)); + } + + // Querying: rows may not be visible yet, see "Read-after-write". + try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + + "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", + { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, + ); + for await (const batch of query) { + // A nonempty replay starts at sequence 0; the callback handles + // empty ones. + if (batch.batchSequence === 0n) recentPrices.length = 0; + for (const row of batch.rows()) recentPrices.push(row); + } + await query.completion; + console.log(recentPrices); + } finally { + await lease.close(); + } + } catch (error) { + if (error instanceof QwpPoolResourceError) { + // No replica was reachable within the failover budget. + console.warn("no replica available for queries:", error.cause); + } else if (error instanceof QwpEgressQueryError) { + console.error(`query failed: status=${error.status} ${error.message}`); + } else { + throw error; + } + } +} finally { + await db.close(); +} +``` + +The query can still miss newly acknowledged rows until WAL apply catches up; +use the [Read-after-write](/docs/connect/clients/nodejs/#read-after-write) pattern for a visibility guarantee. +A replayed batch is idempotent only because this example retains the event's +ID and timestamp and enables table-level deduplication. + +## Next steps + +- [Node.js client guide](/docs/connect/clients/nodejs/) for ingestion and queries. +- [Connect string reference](/docs/connect/clients/connect-string/) for the shared keys. +- [Delivery semantics](/docs/concepts/delivery-semantics/) for replay and deduplication. diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 02393e126f..40e37f577e 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -25,11 +25,11 @@ Key capabilities: - **[Querying](#querying)**: SQL with typed bind parameters, results streamed as columnar batches, DDL and DML execution, cancellation, deadlines, and flow control. -- **[One pooled client](#the-connection-pool)**: `connectQwpNodeClient()` +- **[One pooled client](/docs/connect/clients/nodejs-operations/#the-connection-pool)**: `connectQwpNodeClient()` configures ingestion and queries from one `ws::` connect string, then hands out pooled senders (`db.borrowSender()`) and query leases (`db.borrowQuery()`). -- **[Failover](#failover-and-high-availability)**: multi-host endpoint lists, +- **[Failover](/docs/connect/clients/nodejs-operations/#failover-and-high-availability)**: multi-host endpoint lists, automatic reconnect, and replay of unacknowledged rows. Replay is at least once: pair it with table [deduplication](/docs/concepts/deduplication/) for exactly-once ingestion. @@ -37,18 +37,19 @@ Key capabilities: accepting rows while QuestDB is unreachable and survives process restarts. - **[UDP](#fire-and-forget-udp)**: fire-and-forget ingestion for metrics where occasional loss is acceptable. -- **[Error handling](#error-handling)**: typed errors, asynchronous rejection +- **[Error handling](/docs/connect/clients/nodejs-operations/#error-handling)**: typed errors, asynchronous rejection callbacks, and connection events. The Node.js client differs from the other QWP clients in a few places; see - [Differences from other clients](#differences-from-other-clients). + [Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). :::tip Upgrading from 4.x or using ILP Version 5.0.0 adds QWP and changes how the existing `Sender` handles `null` -and `undefined` values; see [Upgrading from 4.x](#upgrading-from-4x). To move -existing ILP code to QWP, see [From ILP to QWP](#from-ilp-to-qwp). The +and `undefined` values; see [Upgrading from 4.x](/docs/connect/clients/nodejs-operations/#upgrading-from-4x). To move +existing ILP code to QWP, see [From ILP to QWP](/docs/connect/clients/nodejs-operations/#from-ilp-to-qwp). The `Sender` class still speaks ILP over HTTP and TCP; for those transports, see -[ILP transports (legacy)](#ilp-transports-legacy) near the end of this page. +[ILP transports (legacy)](/docs/connect/clients/nodejs-operations/#ilp-transports-legacy) +on the operations and reference page. ::: @@ -176,7 +177,7 @@ What happens: 4. The query handle is an async iterable of result batches, and `batch.rows()` yields one array per row. 5. `db.close()` closes both pools; see - [Closing the pooled client](#closing-the-pooled-client). + [Closing the pooled client](/docs/connect/clients/nodejs-operations/#closing-the-pooled-client). Without the `CREATE TABLE`, the first write creates `trades` automatically, with a designated timestamp column named `timestamp`. A `trades` table that @@ -227,10 +228,10 @@ The `QwpClient` handle has five members: | `borrowQuery()` | `Promise` | Lease an exclusive query connection. Its `close()` returns it to the pool. | | `connect()` | `Promise` | Open the pool minimums. Called for you by `connectQwpNodeClient()`. Safe to retry after a failure. | | `metrics` | `QwpClientMetrics` | Pool counters (`total`, `available`, `leased`, `creating`, `waiting`) for senders and queries. | -| `close()` | `Promise` | Close both pools. Resolves even if rows are not acknowledged; see [Closing the pooled client](#closing-the-pooled-client). Idempotent. | +| `close()` | `Promise` | Close both pools. Resolves even if rows are not acknowledged; see [Closing the pooled client](/docs/connect/clients/nodejs-operations/#closing-the-pooled-client). Idempotent. | Share one `QwpClient` across your application and close it at shutdown. See -[The connection pool](#the-connection-pool) for pool sizing and lease rules. +[The connection pool](/docs/connect/clients/nodejs-operations/#the-connection-pool) for pool sizing and lease rules. ### Standalone Sender @@ -346,10 +347,10 @@ A QWP connect string has the form `schema::key=value;key=value;`: To add settings to a connect string that comes from configuration, such as `QDB_CLIENT_CONF`, append only keys that the string does not set already, or -pass the setting as a [typed option](#programmatic-options), which takes +pass the setting as a [typed option](/docs/connect/clients/nodejs-operations/#programmatic-options), which takes precedence without a duplicate-key error. -The Node.js client's parser differs from some other clients in two places: +The Node.js client's parser differs from some other clients in these ways: - `auto_flush_rows` and `auto_flush_interval` take `0`, not `off`, to disable a trigger. `auto_flush=off` disables auto-flushing entirely. @@ -359,71 +360,26 @@ The Node.js client's parser differs from some other clients in two places: For every key and its default, see the [connect string reference](/docs/connect/clients/connect-string/) and the -[configuration reference](#configuration-reference) at the end of this page. +[configuration reference](/docs/connect/clients/nodejs-operations/#configuration-reference). -### Programmatic options +## Ingestion modes {#ingestion-modes} -Callbacks, custom agents, and other settings a string cannot express go in the -second argument, a `QwpNodeClientConfigOptions` object. When the connect string -and typed options set the same option, the typed value wins. Credentials and -TLS are the exception: a typed `webSocket.authorization` header cannot be -combined with `token`, `username`, or `password` in the string, and a typed -`webSocket.agent` cannot be combined with `tls_verify` or `tls_roots`. Both -combinations are rejected. +The storage choice (`sf_dir`) and the first-connection choice +(`initial_connect_retry` or `lazy_connect`) are independent. This page uses +these names for how a sender publishes and retries: -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - // QwpSenderOptions for every pooled sender - sender: { awaitServerAck: true }, - // Ingestion callbacks and replay settings - ingressSession: { - onSenderError: (error) => - console.error("rejected batch", error.category, error.serverMessage), - }, - // Query session defaults - egressSession: { - queryTimeoutMs: 30_000, - cancelDrainTimeoutMs: 5_000, - serverInfoTimeoutMs: 10_000, - }, - // Egress-only routing and compression - egress: { compression: "zstd" }, - // Pool sizes and timeouts - pool: { senderPoolMax: 2, queryPoolMax: 8 }, -}); -await db.close(); -``` - -The typed `egress` section takes `target`, `zone`, `compression`, -`compressionLevel`, and `maxBatchRows`. The other sections are `webSocket` -(connection settings shared by both directions, such as `agent` or -`connectTimeoutMs`) and `storeAndForward`, the journal settings described -under [Store-and-forward](#store-and-forward). Its fields match the connect -string keys: +| Mode | Enabled by | `flush()` resolves when | During an outage | +|---|---|---|---| +| Default memory mode | Neither `sf_dir` nor a background start | The batch is written to the WebSocket, or queued for replay | `flush()`, auto-flushing `at()`, and a borrowed sender's `close()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | +| Background memory mode | No `sf_dir`; `initial_connect_retry=async` or `lazy_connect=on` | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | +| Store-and-forward | `sf_dir`, with either foreground or background startup | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | -| Typed field | Connect string key | -|---|---| -| `directory` | `sf_dir` | -| `maxBytes` | `sf_max_total_bytes` | -| `maxSegmentBytes` | `sf_max_segment_bytes` | -| `durability` | `sf_durability` | -| `checkpointIntervalMs` | `sf_sync_interval_millis` | -| `appendDeadlineMs` | `sf_append_deadline_millis` | -| `drainOrphans` | `drain_orphans` | -| `maxBackgroundDrainers` | `max_background_drainers` | - -The second argument has no field for `sender_id`; set it in the connect -string. - -`Sender.fromConfig()` takes `{ log, agent, qwp }` as its second argument, -where `qwp` has the sections `webSocket`, `session` (the equivalent of -`ingressSession`), `sender`, and `udp`. - -The `reconnect` objects in `ingressSession` and `egressSession` replace the -whole reconnect policy parsed from the connect string; see -[Typed reconnect policy](#typed-reconnect-policy) before you set one. +Background startup retries the first connection indefinitely, with or without +`sf_dir`. With foreground startup, a sender with `sf_dir` must connect first; +subsequent disconnects are retried indefinitely. In the default memory mode, +a running sender stops after the reconnect budget. See +[Starting while QuestDB is down](/docs/connect/clients/nodejs-operations/#starting-while-questdb-is-down) and +[Ingestion reconnect](/docs/connect/clients/nodejs-operations/#ingestion-reconnect) for startup and outage behavior. @@ -483,14 +439,14 @@ combined with `tls_verify` or `tls_roots`. The pooled client reports connection setup failures as the `cause` of a `QwpPoolResourceError`. For the setup deadlines and the errors they produce, -see [Connection timeouts](#connection-timeouts). +see [Connection timeouts](/docs/connect/clients/nodejs-operations/#connection-timeouts). ### Unsupported authentication paths | Path | Status | Workaround | |---|---|---| | OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider and has no callback to refresh a token. | Obtain an access token from your identity provider, pass it as `token=...`, and create a new client before the token expires. See [OpenID Connect](/docs/security/oidc/). | -| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and senders in default memory mode then fail, while senders with `sf_dir` or in background memory mode keep retrying and buffering (see [Connection-level errors](#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | +| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and senders in default memory mode then fail, while senders with `sf_dir` or in background memory mode keep retrying and buffering (see [Connection-level errors](/docs/connect/clients/nodejs-operations/#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | | Mutual TLS (client certificates) | Not supported. QuestDB does not negotiate client certificates. | Use token or basic authentication over `wss`. | | ILP JWK authentication | Not available for QWP. `auth`, `jwk`, `token_x`, and `token_y` are rejected on `ws`/`wss`. | Use token or basic authentication. | @@ -504,292 +460,10 @@ wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN; ``` Add `tls_roots=/path/to/ca.pem;` when the servers use a private CA. See -[Multiple endpoints](#multiple-endpoints) for routing queries to replicas, and -the [full example](#full-example-ingestion-and-querying-with-failover) for a +[Multiple endpoints](/docs/connect/clients/nodejs-operations/#multiple-endpoints) for routing queries to replicas, and +the [full example](/docs/connect/clients/nodejs-operations/#full-example-ingestion-and-querying-with-failover) for a complete program with this configuration. -## The connection pool - -The pooled client keeps two elastic pools: one of senders and one of query -connections. Each pool opens its minimum on `connect()`, grows on demand up to -its maximum, and a housekeeper closes connections that stay idle too long or -exceed their maximum lifetime, never going below the minimum. - -### Borrowing a sender - -A borrowed sender belongs to the borrower until its `close()` returns it: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - for (const [symbol, price] of [ - ["ETH-USD", 2615.54], - ["BTC-USD", 39269.98], - ] as const) { - await sender - .table("trades") - .symbol("symbol", symbol) - .symbol("side", "buy") - .doubleColumn("price", price) - .doubleColumn("amount", 0.1) - .at(Date.now(), "ms"); - } - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence); - } finally { - // Returns the sender to the pool; any pending rows are flushed. - await sender.close(); - } -} finally { - await db.close(); -} -``` - -A long-running producer can keep its borrow for its whole lifetime and call -`flush()` between batches. Size `sender_pool_max` to the number of producers -that hold a sender at the same time. - -`close()` on a borrowed sender flushes its completed rows and returns it to -the pool. The example above waits for the acknowledgement *before* returning -the sender: `close()` itself does not wait for acknowledgements, although in -default memory mode its flush waits for the reconnect during an outage. See -[Closing a borrowed sender](#closing-a-borrowed-sender) for how long that can -take, how to wait for acknowledgements, and what happens when a close fails. - -After `close()`, every method call or property read on that sender object -throws `QwpClientClosedError`. Don't keep references to a returned sender, for -example in callbacks that can run later. - -### Borrowing a query lease - -A query lease runs one query at a time. For concurrent queries, borrow one lease -per query, up to `query_pool_max`: - -```typescript -import { - connectQwpNodeClient, - type QwpQueryLease, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); - -async function countBySymbol(lease: QwpQueryLease, symbol: string) { - const query = await lease.query( - "SELECT count() FROM trades WHERE symbol = $1", - { binds: (binds) => binds.setVarchar(0, symbol) }, - ); - let count = 0n; - for await (const batch of query) count = batch.get(0, 0) as bigint; - await query.completion; - return count; -} - -try { - const [a, b] = await Promise.all([db.borrowQuery(), db.borrowQuery()]); - try { - // Two leases, two WebSockets: the queries run concurrently. - const [eth, btc] = await Promise.all([ - countBySymbol(a, "ETH-USD"), - countBySymbol(b, "BTC-USD"), - ]); - console.log({ eth, btc }); - } finally { - await Promise.all([a.close(), b.close()]); - } -} finally { - await db.close(); -} -``` - -Starting a second query on a lease while one is still active rejects with -`a QWP query is already active on this connection`. Always close a lease in -`finally`: an unreturned lease holds its connection until `db.close()`. - -### Pool settings - -| Key | Default | Purpose | -|---|---|---| -| `sender_pool_min` | `1` | Senders kept open even when idle. `0` lets the pool close them all. | -| `sender_pool_max` | `4` | Maximum senders the pool opens. | -| `query_pool_min` | `1` | Query connections kept open even when idle. | -| `query_pool_max` | `4` | Maximum query connections, which also caps concurrent queries. | -| `acquire_timeout_ms` | `5000` | How long a borrow waits when the pool is at its maximum, before rejecting with `QwpPoolAcquireTimeoutError`. | -| `idle_timeout_ms` | `60000` | Idle time before an excess connection is closed. `0` keeps idle connections. | -| `max_lifetime_ms` | `1800000` | Age at which an idle connection above the pool minimum is closed. Connections kept open by `sender_pool_min` and `query_pool_min` are never recycled, so this does not rotate a pool that is at its minimum. `0` disables it. | -| `housekeeper_interval_ms` | `5000` | How often the housekeeper checks for idle and over-age connections. Minimum `100`. | -| `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | -| `lazy_connect` | `off` | Start without connecting. See below. | - -Pool sizes, acquisition and idle timeouts, lifetime, and housekeeping settings -have typed equivalents in the `pool` section of the second argument -(`senderPoolMin`, `acquireTimeoutMs`, `housekeepingIntervalMs`, and so on). -The other two settings use different locations: - -- `query_close_timeout_ms` maps to `egressSession.cancelDrainTimeoutMs`, not - `pool`. -- Set `lazy_connect=on` in the connect string. When passing a full - `QwpNodeClientOptions` object instead of a string, use top-level - `lazyConnect: true`. It is not supported in `pool` or the second argument. - -When creating a new pooled connection fails, the borrow rejects with -`QwpPoolResourceError`, whose `cause` holds the connection error. - -`borrowSender()` and `borrowQuery()` take no timeout argument. When the pool -is at its maximum, a borrow waits up to `acquire_timeout_ms` for a connection -to be returned. Opening a new connection is bounded by the -[connection timeouts](#connection-timeouts) of each endpoint, and by the -[failover budget](#query-failover) when query retries are on. To enforce a -shorter deadline, such as a request deadline, race the borrow against a timer -and return a lease that arrives late: - -```typescript -import { connectQwpNodeClient, type QwpClient } from "@questdb/nodejs-client"; - -function borrowQueryWithin(db: QwpClient, timeoutMs: number) { - const borrow = db.borrowQuery(); - let timer: ReturnType | undefined; - const deadline = new Promise((_, reject) => { - timer = setTimeout(() => reject(new Error("borrow timed out")), timeoutMs); - }); - return Promise.race([borrow, deadline]) - .catch((error: unknown) => { - // Return a lease that arrives after the deadline. - borrow.then((lease) => lease.close(), () => undefined); - throw error; - }) - .finally(() => clearTimeout(timer)); -} - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await borrowQueryWithin(db, 2_000); - try { - const query = await lease.query("SELECT count() FROM trades"); - for await (const batch of query) console.log(batch.get(0, 0)); - await query.completion; - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -### Starting while QuestDB is down - -`connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set -`lazy_connect=on` to start regardless: senders connect in the background and -buffer rows in memory until QuestDB is reachable. The query pool stays empty -until the first query. - -```typescript -import { - connectQwpNodeClient, - QwpIngressAckTimeoutError, -} from "@questdb/nodejs-client"; - -// Resolves immediately, even if QuestDB is not running yet. -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;lazy_connect=on;", -); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); - const sequence = sender.publishedSequence; - // Keep the process running until QuestDB comes back and acknowledges it. - for (;;) { - try { - await sender.waitForAcknowledged(sequence, 10_000); - break; - } catch (error) { - if (!(error instanceof QwpIngressAckTimeoutError)) throw error; - console.info("still waiting for QuestDB; do not restage the row"); - } - } - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -Rows buffered while QuestDB is down exist only in memory. This example keeps -the client running until the row is acknowledged; if the process exits first, -the unacknowledged row may be lost. See -[Closing the pooled client](#closing-the-pooled-client). -To keep them across a shutdown or restart, add a -[store-and-forward](#store-and-forward) journal with `sf_dir`. Replay from the -journal is at least once, so write to a deduplicated table as described there. - -`lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, -and rejects an explicit conflicting value. Setting `initial_connect_retry=async` -without `lazy_connect` is not enough: the query pool still connects at startup, -so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query -borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. - -A lazy start does not cover two cases. A locked store-and-forward journal -still fails startup; see [Lock recovery](#sf-lock-recovery). And until a -sender has connected once, it cannot check batches against the server's size -limit; see [Batch size limits](#batch-size-limits). - -### Closing the pooled client - -`db.close()` rejects new borrows, then: - -- Cancels active queries and closes every query connection, including leased - ones. -- Closes idle senders. Each publishes its remaining rows and waits up to - `close_flush_timeout_millis` (5 seconds) for QuestDB to acknowledge them. -- Waits for borrowed senders to be returned, until 5 seconds after - `db.close()` was called, or `acquire_timeout_ms` if that is lower. Closing - the idle senders counts toward the same deadline. A sender still borrowed - after that stays open: its owner must `close()` it, and the process stays - alive until then. - -`db.close()` resolves even when an acknowledgement does not arrive in time. -Without `sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they -stay in the journal, and the next sender on the same directory replays them. -The client usually reports the timeout to `ingressSession.onError` as a -non-terminal `QwpIngressAckTimeoutError`, logged as a warning by default, but -the report is best-effort: do not rely on it to detect unacknowledged rows. To -know that QuestDB accepted every row before shutting down, wait for the -acknowledgement before returning each sender (see -[Awaiting acknowledgements](#awaiting-acknowledgements)), or use -[store-and-forward](#store-and-forward). - -## Concurrency - -Node.js runs your code on one thread, but async functions interleave at every -`await`: - -- **`QwpClient`** is safe to share across your whole application. -- **Senders** are not safe for concurrent producers. A row is built across - several calls, so an `await` between `table()` and `at()` lets another task - add columns to the same row. Give each producer its own sender, borrowed from - the pool, and size `sender_pool_max` to match. -- **Query leases** run one query at a time. Borrow one lease per concurrent - query; `query_pool_max` caps concurrent queries. -- **Worker threads** cannot share clients. Create one client per worker, and - give each worker its own `sender_id` when using store-and-forward. - -Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy -work out of them. Row encoding runs on the event loop too, so one process -ingests at most as fast as one CPU core allows. To go faster, split the stream -across worker threads or processes, each with its own client. - ## Data ingestion @@ -797,7 +471,7 @@ across worker threads or processes, each with its own client. ### General usage pattern A sender is not safe for concurrent producers: the row in progress is shared -state, so borrow one sender per producer (see [Concurrency](#concurrency)). +state, so borrow one sender per producer (see [Concurrency](/docs/connect/clients/nodejs-operations/#concurrency)). 1. Borrow a sender with `db.borrowSender()`, or create a [standalone `Sender`](#standalone-sender). @@ -842,7 +516,7 @@ type for an existing column, after `flush()` has resolved: to the rows, wait for the acknowledgement after flushing, with `await sender.waitForAcknowledged(sender.publishedSequence)`. See [Awaiting acknowledgements](#awaiting-acknowledgements) and -[Ingestion errors](#ingestion-errors). +[Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). Tables and columns are created automatically, with the column types listed below. Table and column names are validated locally with QuestDB's rules @@ -859,7 +533,7 @@ drops every row staged since the last flush. An awaited `at()` or `atNow()` can also reject because an auto-flush failed after the row was completed. Whether the completed rows are still staged, and what to do next, depends on the error class; see the -[Error handling](#error-handling) table. +[Error handling](/docs/connect/clients/nodejs-operations/#error-handling) table. ### Column methods @@ -971,7 +645,7 @@ names that differ only in case. For example, not raise a type mismatch. Invalid values can still fail local validation. For an existing table, QuestDB rejects an incompatible type or value -asynchronously; see [Ingestion errors](#ingestion-errors). Compatible +asynchronously; see [Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). Compatible conversions are allowed: for example, `longColumn("price", 123n)` can write to an existing DOUBLE column. This does not change the sender's local type-consistency rule. @@ -1131,7 +805,7 @@ have 1 to 32 dimensions. Only DOUBLE arrays can be ingested: `longArrayColumn()` exists for protocol parity, but current servers reject it with `long arrays are not supported, only double arrays`. The rejection is terminal; see -[Recovering from a terminal rejection](#recovering-from-a-terminal-rejection). +[Recovering from a terminal rejection](/docs/connect/clients/nodejs-operations/#recovering-from-a-terminal-rejection). Query results return arrays as `{ dimensions, values }`; see [Reading result values](#reading-result-values). @@ -1182,9 +856,10 @@ try { ``` - `decimalColumnText()` takes a - plain decimal string (such as `"0.0750"`) and preserves the literal's scale, - including trailing zeros. Scientific notation is accepted for a `number`, - not a string; JavaScript drops trailing zeros when formatting numbers. + decimal string (such as `"0.0750"`) and preserves the literal's scale, + including trailing zeros. Strings and numbers both accept scientific notation + (such as `"1.5e-3"`); pass a string when scale matters, because JavaScript + drops trailing zeros when formatting numbers. - `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` or as big-endian two's-complement bytes in an `Int8Array`. @@ -1318,15 +993,7 @@ clamped to 90% of the effective batch limit: the server's limit, or `sf_max_segment_bytes` when that is lower (see [Batch size limits](#batch-size-limits)). -What `flush()` waits for depends on the ingestion mode. This page uses these -three names for the modes: - -| Mode | Enabled by | `flush()` resolves when | During an outage | -|---|---|---|---| -| Default memory mode | Neither of the others | The batch is written to the WebSocket, or queued for replay | `flush()`, auto-flushing `at()`, and a borrowed sender's `close()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | -| Background memory mode | `initial_connect_retry=async` or `lazy_connect=on` | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | -| Store-and-forward | `sf_dir` | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | - +The [ingestion mode](#ingestion-modes) determines when `flush()` resolves. In every mode, `flush()` does not wait for QuestDB to acknowledge the rows, unless you set `awaitServerAck`. Unacknowledged batches are kept and replayed after a reconnect. See [Awaiting acknowledgements](#awaiting-acknowledgements) @@ -1374,8 +1041,9 @@ Call `reset()` to drop every row staged since the last flush, then write the rows again without the oversized one. Until a sender has connected once, it does not know the server's limit. This -applies in background memory mode, and to a store-and-forward sender that -restarts while QuestDB is down. Batches are then capped only by +applies in background memory mode and to a store-and-forward sender with +`lazy_connect=on` or `initial_connect_retry=async` that starts while QuestDB +is down. Batches are then capped only by `sf_max_segment_bytes`: 4 MiB with `sf_dir`, and no cap without it. A batch larger than the server's limit passes `flush()` but can never be delivered: the sender keeps reconnecting, and `waitForAcknowledged()` times out. With @@ -1457,7 +1125,7 @@ asynchronously, a sender can fail after its `close()` already succeeded: the error then surfaces on the next borrower's auto-flushing `at()`, `flush()`, or `close()`, and the pool replaces the sender after that. The next borrower's own staged rows are lost with the failed sender, even rows for other tables: write -them again on a new borrow. See [Ingestion errors](#ingestion-errors). +them again on a new borrow. See [Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). ### Awaiting acknowledgements @@ -1528,7 +1196,7 @@ To make every `flush()` wait for its acknowledgement, set `awaitServerAck`: `connectQwpNodeClient(conf, { sender: { awaitServerAck: true } })`, or `{ qwp: { sender: { awaitServerAck: true } } }` for a standalone `Sender`. A server rejection then rejects the waiting `flush()` itself. See -[Ingestion errors](#ingestion-errors) for the error classes before and after +[Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors) for the error classes before and after a terminal failure. Acknowledgement is not required for delivery: unacknowledged batches are @@ -1759,10 +1427,11 @@ try { ``` With a journal, the sender keeps accepting rows while QuestDB is unreachable, -subject to [journal capacity](#sf-capacity). It retries the connection -indefinitely once it has connected, and a new sender opened on the same -directory replays what the previous process left behind, once it can take over -the directory's lock (see [Lock recovery](#sf-lock-recovery)). +subject to [journal capacity](#sf-capacity). With `lazy_connect=on` as above, +it retries from startup; without background startup, the first connection +must succeed, then later disconnects are retried indefinitely. A new sender +opened on the same directory replays what the previous process left behind, +once it can take over the directory's lock (see [Lock recovery](#sf-lock-recovery)). - **Layout.** A standalone `Sender` journals into `/`. A pooled client uses one directory per pooled sender: @@ -1782,14 +1451,14 @@ the directory's lock (see [Lock recovery](#sf-lock-recovery)). disk sync to every flush: on a producer that flushes often, prefer `periodic` or larger batches. - **Startup.** To start the pooled client while QuestDB is down, see - [Starting while QuestDB is down](#starting-while-questdb-is-down). A + [Starting while QuestDB is down](/docs/connect/clients/nodejs-operations/#starting-while-questdb-is-down). A standalone `Sender` needs only `initial_connect_retry=async` or `lazy_connect=on`. With the default `initial_connect_retry=off`, the first connection must succeed. - **Rejected batches.** A batch that QuestDB rejects terminally stays at the head of the journal and stops ingestion through the client, for every table, until you act; see - [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection). + [Recovering from a terminal rejection](/docs/connect/clients/nodejs-operations/#recovering-from-a-terminal-rejection). - **Orphans.** With `drain_orphans=on`, a sender also adopts and drains journals with other `sender_id` values left under the same `sf_dir` by processes that crashed, up to `max_background_drainers` (4) at a time. @@ -1897,16 +1566,17 @@ try { } ``` -If the server does not support durable acknowledgement, connecting fails with +If the server does not support durable acknowledgement, a sender that connects +in the foreground before its first successful connection fails with `QwpDurableAckUnavailableError`, which the pooled client reports as the -`cause` of a `QwpPoolResourceError`. - -Senders in background memory mode instead keep retrying and emit -`durable-ack-unavailable` -[connection events](#connection-events). A store-and-forward sender does the -same when reconnecting after its first successful connection. Monitor these -events and buffer usage: successful background startup does not confirm that -the server supports durable acknowledgement. +`cause` of a `QwpPoolResourceError`. A background-started sender +(`initial_connect_retry=async` or `lazy_connect=on`) instead retries from +startup and emits `durable-ack-unavailable` +[connection events](/docs/connect/clients/nodejs-operations/#connection-events), **even with `sf_dir`**. With `sf_dir` +and a foreground start, the first connection fails, but a sender that has +connected successfully before keeps retrying after a later mismatch. Monitor +these events and buffer usage: successful background startup does not confirm +that the server supports durable acknowledgement. ### Fire-and-forget UDP @@ -2000,7 +1670,7 @@ try { | `autoCredit` | `true` | Replenish the credit window as batches are consumed. | | `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | -There is no per-query failover setting; see [Query failover](#query-failover). +There is no per-query failover setting; see [Query failover](/docs/connect/clients/nodejs-operations/#query-failover). The `QwpEgressQuery` handle has these members: | Member | Purpose | @@ -2037,7 +1707,7 @@ query again from its first batch. If your loop accumulates rows, reset them when `batch.batchSequence === 0n`. A replay that returns **no batches** has no sequence to detect, so also clear accumulated state on `onReplayReset` (on a client with only one active query), or use `failover=off` and retry the whole -query. See [Query failover](#query-failover). +query. See [Query failover](/docs/connect/clients/nodejs-operations/#query-failover). ### Reading result values @@ -2556,7 +2226,7 @@ try { For row-major work instead, `batch.forEachRow()` reuses one row object; read values inside its callback only when they contribute to your result. For an accumulator that supports automatic query replay, see -[Query failover](#query-failover). +[Query failover](/docs/connect/clients/nodejs-operations/#query-failover). Batches are delivered one at a time: when the callback returns a promise, the client waits for it before delivering the next batch. With a credit window set @@ -2590,1049 +2260,17 @@ only. `zoneId`, `clusterId`, `nodeId`, and `capabilities`. It refreshes after a failover. -## Error handling - -Each error leaves the client in a known state. The sections after this table -have the details and examples: - -| Error | Surfaces from | State afterwards | What to do | -|---|---|---|---| -| `TypeError`, `RangeError`, or `Error` from local validation | The column method or `at()` that staged the value | The row in progress is discarded; the sender stays usable | Fix the value and write the row again | -| `QwpBatchTooLargeError` | `flush()`, the `at()` whose auto-flush sends the batch, or `close()` | The batch can never be sent: the staged rows are kept, every later flush fails the same way, and `close()` discards them | Call `reset()`, then write the rows again without the oversized one; see [Batch size limits](#batch-size-limits) | -| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, an auto-flushing `at()`, or `close()` | The batch stays staged; the sender stays usable | Keep the sender and flush again later. Don't write the rows again, and don't close the sender while backpressure persists; see [Backpressure](#backpressure) | -| Retriable server rejection | `onSenderError` | The client resends the batch; repeated rejections become terminal | Monitor; no action needed per rejection | -| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection) | -| `QwpReconnectExhaustedError` on a sender | `onError` with `terminal: true`, then the next `flush()`, `at()`, or `close()` | The sender has failed; unsent rows are lost | Borrow a new sender; see [Ingestion reconnect](#ingestion-reconnect) | -| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See [Lock recovery](#sf-lock-recovery) | -| `QwpPoolResourceError` with another `cause` | `connectQwpNodeClient()`, `borrowSender()`, or `borrowQuery()` | No connection was opened | Unwrap `cause`; see [Connection-level errors](#connection-level-errors) | -| `QwpEgressQueryError` | Query iteration and `completion` | The lease stays usable | Fix the SQL or the bind values | -| `QwpEgressQueryTimeoutError`, `QwpEgressQueryAbandonedError` | Query iteration and `completion` | The lease is busy until QuestDB confirms the cancellation | Close the lease and borrow a new one | -| `QwpEgressQueryCancelTimeoutError` | Query iteration and `completion` | The connection is closed | Close the lease and borrow a new one | -| `QwpReconnectExhaustedError` on a query | Query iteration and `completion` | The lease stays failed, even after QuestDB recovers | Close the lease and borrow a new one; see [Query failover](#query-failover) | - -`QwpReplayStoreAppendTimeoutError` and the other store-and-forward journal -errors extend `QwpReplayStoreError`, so test for the specific classes before -the base class, or branch on `error.retryable`. `true` means the failure is -temporary and the sender stays usable. `false` means the journal itself can no -longer be used, for example `QwpReplayStoreCorruptionError` or -`QwpReplayStoreLockLostError`, and the sender has failed. The other classes in -the table have no common base class: test each with `instanceof`. - -### Ingestion errors - -Ingestion reports errors in two ways: - -- **While building a row.** A column method throws, or the promise returned by - `at()` or `atNow()` rejects, with a `TypeError`, `RangeError`, or `Error` for - an invalid value or name. The row in progress is discarded, and the sender - stays usable. -- **Asynchronously, when QuestDB rejects a batch.** The rejection arrives after - `flush()` resolved. It is delivered to the `onSenderError` callback, and - surfaces as a rejection of `waitForAcknowledged()`, or of `flush()` with - `awaitServerAck`. - -```typescript -import { - connectQwpNodeClient, - QWP_SENDER_ERROR_POLICY, - type QwpSenderError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - ingressSession: { - onSenderError: (error: QwpSenderError) => { - // serverStatusByte is absent for client-side errors. - const status = - error.serverStatusByte === undefined - ? "none" - : `0x${error.serverStatusByte.toString(16)}`; - console.error( - `rejected [${error.category}, policy=${error.appliedPolicy}, ` + - `status=${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + - error.serverMessage, - ); - if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { - // The sender stopped: alert, and fix the data or the schema. - } - }, - }, -}); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence, 10_000); - } finally { - // Rethrows a terminal error. The pool then replaces the sender. - await sender.close(); - } -} finally { - await db.close(); -} -``` - -When `onSenderError` is not set, rejections are logged: retriable ones at -`warn`, terminal ones at `error`. Callbacks run asynchronously, never inside the -client's protocol handling, and an exception thrown by a callback is contained. -A standalone sender's `close()` can also reject, with -`QwpSenderCloseTimeoutError`, when its rows are not acknowledged in time; see -[Closing a sender](#closing-a-sender). - -`QwpSenderError` fields: - -| Field | Type | Meaning | -|---|---|---| -| `category` | `string` | `schema-mismatch`, `parse-error`, `security-error`, `write-error`, `internal-error`, `not-writable`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `protocol-violation`, `data-loss`, or `unknown`. Branch on this field. | -| `appliedPolicy` | `string` | What the client did: `retriable` (reconnect and resend), `retriable-other` (resend to another endpoint), `terminal` (the sender stopped), or `abandoned` (journaled data was quarantined). | -| `serverStatusByte` | `number` | The raw QWP status code, for example `0x03` for a schema mismatch. Absent for client-side errors. | -| `serverMessage` | `string` | QuestDB's error text, for example `cannot parse DOUBLE from string [value=abc, column=price]`. | -| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `publishedSequence`. | -| `messageSequence` | `bigint` | The wire sequence of the rejected message. | -| `tableName` | `string` | The table, when the server attributes the rejection to one. Often absent. | -| `detectedAtMs` | `number` | When the client received the rejection. | -| `quarantinedPath` | `string` | For `data-loss` in store-and-forward: where the unreplayable journal was preserved. | - -The default policy follows the category: - -| Category | Policy | Examples | -|---|---|---| -| `schema-mismatch`, `parse-error`, `security-error`, `protocol-violation` | Terminal | Wrong value type for an existing column, malformed data, missing permission | -| `write-error`, `internal-error`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `unknown` | Retriable | Disk pressure, a suspended table, a transient server fault | -| `not-writable` | Retriable on another endpoint | The server is a replica or cannot accept writes | -| `data-loss` | Abandoned | A corrupt store-and-forward journal was set aside | - -A retriable rejection is resent. For rejections that count toward the -poison-frame detector, if the same batch keeps being rejected after -`max_frame_rejections` (4) attempts spanning at least -`poison_min_escalation_window_millis` (5 minutes), the sender stops as for a -terminal error. The `dictionary-gap`, `unknown`, and `not-writable` categories -are exempt: they reset the poison episode instead of adding a strike. -Retriable rejections of symbol-dictionary catch-up frames are also exempt. -The six `on_*_error` connect-string keys are accepted but not applied by this -client. - -Handling notes: - -- **Message stability.** `serverMessage` is free-form English text from the - server. Its wording can change between releases: branch on `category`, not on - the text. -- **Sensitive data.** Server messages can contain column names and values. - Treat them as untrusted input, and redact them before sending them to - third-party error trackers or showing them to end users. -- **Correlation.** There is no server-side request ID. Correlate with the frame - sequence range, `tableName`, and `detectedAtMs`. - -#### Recovering from a terminal rejection - -After a terminal server rejection, the sender is permanently failed. An -already-pending `waitForAcknowledged()` for the rejected batch can reject with -`QwpIngressNackError`. Once the terminal failure is latched, new calls to -`waitForAcknowledged()`, `flush()`, or `close()` reject with -`QwpReplayRejectedError`, whose `status` and message repeat the server's. -Error handlers must allow either class depending on timing. Writing LONG arrays -with `longArrayColumn()` triggers a terminal rejection on every current server. - -Close the sender and create a new one. A pooled sender is replaced -automatically after the `close()` that reports the error. What happens to the -rejected batch depends on the mode: - -- **Without store-and-forward**, the failed sender's unacknowledged batches, - including the rejected one, are discarded with it, and so are rows that a - later borrower staged on it before the error surfaced. The new sender starts - empty. -- **With store-and-forward**, the rejected batch stays at the head of the - journal. Every new sender on that directory, including the pool's - replacement sender and the same client after a restart, sends it again and - fails the same way, with `QwpReplayRejectedError`. Pooled borrows keep - getting that journal, so ingestion through the client stops for every table, - not only the table in the rejected batch, until you act. Treat it as an - outage and alert on it from `onSenderError`. Fix the cause so that QuestDB - accepts the batch, for example by adjusting the table schema, or stop the - process and move the journal directory aside. Moving it aside discards every - unacknowledged batch in it, not only the rejected one. - -### Query errors - -Query errors reject both the `for await` iteration and `completion`: - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, - QwpEgressQueryTimeoutError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM no_such_table"); - for await (const batch of query) console.log(batch.rowCount); - await query.completion; - } catch (error) { - if (error instanceof QwpEgressQueryError) { - // Prints: 5 [14] table does not exist [table=no_such_table] - console.error(error.status, error.message); - } else if (error instanceof QwpEgressQueryTimeoutError) { - console.error("timed out"); - } else { - throw error; - } - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -`QwpEgressQueryError` has `status` (the QWP status code), `message` (the server -text, where parse errors start with the character position in brackets), and -`requestId` (a client-assigned `bigint` that numbers the queries of a -connection). The lease remains usable after a `QwpEgressQueryError`. - -| Status | Name | Meaning | -|---|---|---| -| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column, or a bind value the statement cannot use, such as a boolean for `LIMIT` | -| `0x06` | INTERNAL_ERROR | Execution failure, including a bind value that cannot be converted, such as `'abc'` compared with a DOUBLE column, and a column type that QWP cannot return | -| `0x08` | SECURITY_ERROR | Missing permission | -| `0x0a` | CANCELLED | The query was cancelled with `cancel()` | -| `0x0b` | LIMIT_EXCEEDED | A server limit was reached: the server-side query timeout, memory, or a result row too large to send | - -The `QWP_STATUS` export names these codes, for example -`QWP_STATUS.PARSE_ERROR`, so code can compare against constants instead of -numbers. A status alone does not separate a client mistake from a server -fault: `0x06` covers both bind values that cannot be converted and execution -failures, and `0x0b` covers both the server-side query timeout and memory -limits. - -Other query errors: - -| Error | Meaning | -|---|---| -| `QwpEgressQueryTimeoutError` | The query deadline expired and cancellation started. Has `requestId` and `timeoutMs`. | -| `QwpEgressQueryAbandonedError` | Iteration ended early, for example with `break`. | -| `QwpEgressQueryCancelTimeoutError` | QuestDB did not confirm a cancellation in time; the connection was closed. | -| `QwpEgressSessionClosedError` | The query connection is closed. | -| `QwpReconnectExhaustedError` | Failover gave up; see [Query failover](#query-failover). | - -As with ingestion, the message text is not stable, may echo parts of the SQL, -and has no server-side correlation ID beyond `requestId`. - -### Connection-level errors - -| Error | Raised when | -|---|---| -| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, or `transport`. It also carries `statusCode`, `retryable`, and `url`. | -| `QwpFailoverError` | Every endpoint in a multi-host list failed. `attempts` holds each endpoint and its error. | -| `QwpPoolResourceError` | The pool could not open a new connection. `cause` holds the error above. | -| `QwpPoolAcquireTimeoutError` | Every pooled connection stayed leased past `acquire_timeout_ms`. | -| `QwpReconnectExhaustedError` | The reconnect budget ran out. The sender or query failed permanently. | -| `QwpRoleMismatchError` | No endpoint has the role that `target` requires. | -| `QwpDurableAckUnavailableError` | `request_durable_ack=on`, but the server does not support it. | -| `QwpClientClosedError` | The pooled client, or a returned lease, is already closed. | - -The pooled client wraps every failure to open a connection, from -`connectQwpNodeClient()`, `db.connect()`, `borrowSender()`, or -`borrowQuery()`, in a `QwpPoolResourceError`. Unwrap its `cause` before -checking for a specific error. When `addr` lists several hosts, the cause is a -`QwpFailoverError` whose `attempts` hold the error of each endpoint. When -initial-connect retry is on, for example with a `failover_*` or `reconnect_*` -key, or with a typed `egressSession.reconnect` object for query connections -(see [Typed reconnect policy](#typed-reconnect-policy)), the cause is a -`QwpReconnectExhaustedError` instead, and its own `cause` holds the last -attempt's error: - -```typescript -import { - connectQwpNodeClient, - QwpFailoverError, - QwpPoolResourceError, - QwpReconnectExhaustedError, - QwpUpgradeError, -} from "@questdb/nodejs-client"; - -// The errors behind a failed connection, one per endpoint tried. -function connectionErrors(error: unknown): unknown[] { - let cause = error instanceof QwpPoolResourceError ? error.cause : error; - // With initial-connect retry on, the last attempt's error is wrapped. - if (cause instanceof QwpReconnectExhaustedError) cause = cause.cause; - return cause instanceof QwpFailoverError - ? cause.attempts.map((attempt) => attempt.error) - : [cause]; -} - -try { - const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); - await db.close(); -} catch (error) { - for (const cause of connectionErrors(error)) { - if (cause instanceof QwpUpgradeError && cause.kind === "authentication") { - console.error("QuestDB rejected the credentials:", cause.message); - } else { - console.error("cannot connect:", cause); - } - } - throw error; -} -``` - -An authentication rejection (HTTP 401 or 403) is terminal before a sender's -first successful connection and for query connections. It stops the endpoint -walk because credentials are assumed to be shared across the cluster. - -After a successful connection, regular senders with `sf_dir` or in background -memory mode (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying -authentication rejections indefinitely. This lets buffered data drain once -server-side authentication is restored. Senders in default memory mode, and -orphan drainers, do not have this exception. See -[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) -for how other clients behave. - -Endpoints in error messages have any embedded credentials removed. - -#### Connection timeouts - -Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS -and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and -authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits -`connect_timeout` when only the latter is set. A timeout in either phase -produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or -`authentication`. - -After the upgrade, a query connection has a separate 5-second deadline for -the initial QWP `SERVER_INFO` frame. Configure it with the typed option -`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not -change it. Expiry produces an ordinary `Error` with the message -`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. - -### Logging - -The client writes its own messages to the console by default, at the `error`, -`warn`, and `info` levels. To route a sender's messages, such as warnings about -rows discarded on close, through your logger, pass a `QwpSenderLogger` -function: `{ sender: { log } }` as the second argument of -`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. Its signature -is `(level: "error" | "warn" | "info" | "debug", message: string | Error)`, so -convert `message` with `String()` if your logger takes strings only. The -function also receives `debug` messages, one per staged row, so filter by -level. - -Rejected batches and session errors go to `ingressSession.onSenderError` and -`ingressSession.onError`. Their defaults log to the console, so replace both to -route them through your logger. Some messages from other parts of the client, -such as store-and-forward recovery, always go to the console. - -## Failover and high availability - -:::note Enterprise - -Failing over between several QuestDB hosts requires QuestDB Enterprise -replication. Reconnecting to a single restarted server works in open source -too. - -::: - -### Multiple endpoints - -List several hosts in `addr`: - -```text -wss::addr=db-a.example.com:9000,db-b.example.com:9000,db-c.example.com:9000; -``` - -The client ranks endpoints by observed health and by `zone`, and on a -connection loss moves to the next usable one. `addr` is shared by ingestion and -queries. - -Ingestion always needs the primary: replicas refuse writes, and the sender -walks the list until it finds the current primary. Queries can use any node. -`target` selects which roles queries accept: `any` (the default), `primary`, or -`replica`. Set it with the typed `egress` option, as below, because in the -connect string `target` also filters ingestion (see the caution that follows). - -`target` is a strict filter, not a preference. With `replica`, queries never -fall back to the primary, and they fail when no replica is reachable, including -against a single open source server. Because the pooled client opens a query -connection at startup, `connectQwpNodeClient()` then fails too, with a -`QwpPoolResourceError` whose `cause` leads to a `QwpRoleMismatchError`; see -[Connection-level errors](#connection-level-errors) to unwrap it. To start -without a replica, also set `query_pool_min=0`. Queries borrowed before a -replica is reachable then reject with `QwpPoolResourceError`: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -// Queries run on replicas only. Ingestion still follows the primary. -const db = await connectQwpNodeClient( - "wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN;" + - // Start, and ingest, even while no replica is reachable. - "query_pool_min=0;", - { egress: { target: "replica" } }, -); -await db.close(); -``` - -`zone` prefers endpoints in the same zone. - -:::caution `target` in the connect string also filters ingestion - -Unlike the Java client, the Node.js client applies `target` and `zone` from -the connect string to ingestion as well as queries. `target=replica` in the -connect string therefore stops ingestion from reaching the primary. To read -from replicas and write to the primary with one client, keep `target` out of -the connect string and set it for queries only: -`connectQwpNodeClient(conf, { egress: { target: "replica" } })`. - -::: - -### Ingestion reconnect - -When the connection drops, the sender reconnects with exponential backoff and -jitter, then resends every unacknowledged batch: - -| Key | Default | Purpose | -|---|---|---| -| `reconnect_initial_backoff_millis` | `100` | First retry delay. | -| `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | -| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in default memory mode. `0` removes the limit. | -| `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | - -Whether the sender gives up depends on the mode (see [Flushing](#flushing)): - -- **Default memory mode** retries for up to `reconnect_max_duration_millis` - per outage. When the budget runs out, the sender fails permanently with - `QwpReconnectExhaustedError`, and its unsent rows are lost. The Java - reference client retries indefinitely in this mode instead. -- **Background memory mode** (`initial_connect_retry=async` or - `lazy_connect=on`) and **store-and-forward** (`sf_dir`) retry indefinitely. - -Setting any `reconnect_*` key also makes a sender's first connection retry -within the budget, as if `initial_connect_retry=on`. Set -`initial_connect_retry=off` explicitly to keep a fail-fast start. The keys do -not apply to query connections: the pooled client still opens its query pool -at startup, so `connectQwpNodeClient()` fails fast while QuestDB is down unless -you also set `query_pool_min=0` or enable query retries (see -[Connection events](#connection-events)). - -Replay after a reconnect is at least once: a batch that QuestDB committed just -before the connection dropped is sent again. Write to a deduplicated table, as -described under [Store-and-forward](#store-and-forward), to keep replayed rows -from inserting duplicates. - -### Query failover - -If the connection fails during a query, the client reconnects, to another -endpoint when there is one, and runs the query again from the start: - -| Key | Default | Purpose | -|---|---|---| -| `failover` | `on` | Set `off` to fail the query instead of retrying. | -| `failover_max_attempts` | `8` | Connection attempts per failure. Each attempt tries every endpoint in `addr`. | -| `failover_backoff_initial_ms` | `50` | First retry delay. | -| `failover_backoff_max_ms` | `1000` | Longest delay between retries. | -| `failover_max_duration_ms` | `30000` | Time budget per failure. | - -The attempt limit and the time budget apply together, and whichever is reached -first ends the failover. When attempts fail fast, for example with connection -refused while a server restarts, the 8 attempts and their backoff of 50 ms to -1 second, with jitter, take only about 1 to 3 seconds, and at most about 4.5 -seconds, long before the 30-second budget. To ride out a longer restart, raise `failover_max_attempts`, -or set `maxAttempts: 0` in a typed `egressSession.reconnect` object to remove -the attempt limit and rely on the time budget alone; see -[Typed reconnect policy](#typed-reconnect-policy). - -When failover gives up, the query rejects with `QwpReconnectExhaustedError`, -and the lease stays failed even after QuestDB recovers: close it and borrow a -new one. A `QwpEgressQueryError` from the server is a query result and never -triggers failover. Replaying an in-flight `query()` also re-executes DDL and -DML: an `INSERT` may run twice if its completion was lost. For non-idempotent -SQL, use a separate client configured with `failover=off` and check an -uncertain outcome before retrying; see -[DDL and DML statements](#ddl-and-dml-statements). - -:::warning Clear partial results when a query restarts - -A re-executed query starts again from the first row. Batches that were queued -but not yet consumed are discarded for you, but rows your loop already -processed are delivered again. If your code accumulates rows, clear them when -the query restarts; otherwise it sees the first part of the result twice. - -::: - -Every batch has a `batchSequence` starting at `0n`, including the first batch -after a replay. Clear accumulated rows on that batch. A replay may return -**zero rows and no batches**, though, leaving prior rows in your accumulator. -`egressSession.onReplayReset` also clears it when that happens. Since this -callback is shared across the pool and request IDs are per connection, the -example limits the pool to one active query: - -```typescript -import { - connectQwpNodeClient, - QwpReconnectExhaustedError, -} from "@questdb/nodejs-client"; - -const rows: (readonly unknown[])[] = []; -const db = await connectQwpNodeClient( - "ws::addr=db-a.example.com:9000,db-b.example.com:9000;query_pool_max=1;", - { - egressSession: { - onReplayReset: () => { - rows.length = 0; - }, - }, - }, -); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM trades LIMIT 100000", { - // The deadline covers the whole query, including a re-execution. - timeoutMs: 30_000, - }); - for await (const batch of query) { - // Sequence 0 starts the result, both initially and after a failover. - if (batch.batchSequence === 0n) rows.length = 0; - for (const row of batch.rows()) rows.push(row); - } - await query.completion; - console.log(`${rows.length} rows`); - } catch (error) { - if (!(error instanceof QwpReconnectExhaustedError)) throw error; - // Failover gave up. This lease stays failed: return it, retry later. - console.error("no endpoint could run the query:", error.message); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -The reset needs the rows of the current attempt in one place you can discard. -When your code passes rows on as they arrive, for example by streaming them to -an HTTP response, it cannot take them back after a restart. Then choose one -of these instead: - -- Run such queries on a separate client with `failover=off`, so that a lost - connection fails the query instead of restarting it, and retry the whole - request. -- Keep each result small enough to buffer, for example by paging with - `WHERE timestamp < $1 ORDER BY timestamp DESC LIMIT n`, binding the oldest - timestamp of the previous page. -- Treat `batchSequence === 0n` after rows have left as an error, and abort the - downstream response instead of sending duplicates. - -`egressSession.onReplayReset` runs before a query is replayed, including when -that replay returns no batches. Its event has `requestId`, `endpoint`, -`previousEndpoint`, `serverInfo`, and `cause`. The `requestId` matches -`query.requestId`, but request IDs are numbered per connection and every lease -of a pooled client shares the callback: it cannot identify which of several -concurrent queries restarted. Use a dedicated, single-query client when the -callback resets result state, as above. For concurrent results that cannot be -isolated, set `failover=off` and retry the whole query after a transport error; -`batchSequence === 0n` alone cannot detect a zero-batch replay. - -### Typed reconnect policy - -Reconnect and failover behavior comes from the connect-string keys above, or -from typed `reconnect` objects in the second argument of -`connectQwpNodeClient()`: `ingressSession.reconnect` for senders and -`egressSession.reconnect` for queries. You need the object to register -`onEvent` for [connection events](#connection-events). - -:::caution A typed `reconnect` object replaces the connect-string keys - -`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces -the whole reconnect policy parsed from the `reconnect_*`, -`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and -`egressSession.reconnect` replaces the policy parsed from `failover*` keys. -Fields you leave out of the object take the built-in defaults, not the values -from the connect string. When you supply the object, for example to register -`onEvent`, set every bound you rely on in it. - -::: - -The object's fields and the connect-string keys they replace: - -| Field | Ingestion key, default | Query key, default | -|---|---|---| -| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8`. The key accepts `1` or more; the typed field also accepts `0`, unlimited | -| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | -| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | -| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | -| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | -| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | -| `onEvent` | None | None | - -`egressSession: { reconnect: false }` is the typed equivalent of -`failover=off`. Senders in background memory mode or with `sf_dir` retry -indefinitely: `maxAttempts` and `maxDurationMs` do not end their retries, so -the full example's `reconnect: { onEvent }` with `sf_dir` keeps retrying -through an outage of any length. - -The two directions treat the first connection differently: - -- **Senders**: setting any `reconnect_*` key makes the first connection retry - within the budget, as if `initial_connect_retry=on`. A typed - `ingressSession.reconnect` object does not, so the first connection still - fails fast. Set `initial_connect_retry` in the connect string to choose the - startup behavior. -- **Query connections**: the first connection retries within the failover - budget, for retryable errors, when you supply an `egressSession.reconnect` - object, set `failover=on` explicitly, or set a `failover_*` key without - `failover=off`. Otherwise it is attempted once. `failover=off` and - `egressSession.reconnect: false` turn reconnects off entirely. - -### Connection events - -Register `reconnect.onEvent` to observe connections. Events are delivered -asynchronously through a bounded queue (64 by default, -`connection_listener_inbox_capacity`); when it overflows, the oldest events are -dropped and counted in the metrics. - -```typescript -import { - connectQwpNodeClient, - QWP_RECONNECT_EVENT_KIND, - type QwpReconnectEvent, -} from "@questdb/nodejs-client"; - -function onEvent(event: QwpReconnectEvent) { - switch (event.kind) { - case QWP_RECONNECT_EVENT_KIND.RECONNECTING: - console.warn("connection lost, reconnecting:", event.cause); - break; - case QWP_RECONNECT_EVENT_KIND.FAILED_OVER: - console.warn(`failed over to ${String(event.endpoint)}`); - break; - default: - console.info(event.kind, String(event.endpoint ?? "")); - } -} - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - // Each object replaces the reconnect_* or failover* keys from the connect - // string. Fields left out use the built-in defaults. - ingressSession: { - reconnect: { onEvent }, - // No event marks a terminal failure: it arrives here instead. - onError: (event) => { - if (event.terminal) console.error("ingestion stopped:", event.error); - }, - }, - egressSession: { reconnect: { onEvent } }, -}); -await db.close(); -``` - -Supplying the `reconnect` objects also changes how the first connection is -retried; see [Typed reconnect policy](#typed-reconnect-policy). - -| Kind | Meaning | -|---|---| -| `connected` | The first connection succeeded. | -| `reconnecting` | The active connection was lost. `cause` holds the error. | -| `attempt-failed` | One connection attempt failed. The client retries if the error is retryable and its budget allows; otherwise this is the last event before the failure is reported. | -| `reconnected` | Reconnected to the same endpoint. | -| `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | -| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. Only senders that retry indefinitely wait: store-and-forward senders after their first connection, and senders in background memory mode. | -| `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | -| `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | - -The `QWP_RECONNECT_EVENT_KIND` constants name these kinds: `CONNECTED`, -`RECONNECTING`, `ATTEMPT_FAILED`, `RECONNECTED`, `FAILED_OVER`, -`DURABLE_ACK_UNAVAILABLE`, `DURABLE_ACK_PERSISTENT_FAILURE`, and -`PRIMARY_UNAVAILABLE`. - -`reconnected` and `failed-over` are mutually exclusive: code that tracks the -current node must handle both. The client has no property that says whether it -is connected right now. To report it, for example in a health check, track the -latest event: after `reconnecting` the connection is down, and `connected`, -`reconnected`, or `failed-over` mean it is up. A query that reconnects runs -again from its first batch; see -[Query failover](#query-failover) for resetting accumulated rows. - -No event marks a terminal failure. When a sender stops retrying, because its -reconnect budget ran out or the error cannot be retried, -`ingressSession.onError` receives a `QwpIngressErrorEvent` with -`terminal: true`, even while the sender is idle. The event also has `error`, -`timestampMs`, and, for a server rejection, `senderError`. The sender's next `flush()`, auto-flushing `at()`, or -`close()` then rejects with the same error, such as -`QwpReconnectExhaustedError`. A query that cannot fail over rejects its -iteration and `completion` instead. - -For ingestion, `ingressSession` also accepts `onProgress`, for published, -acknowledged, and durably acknowledged sequences, and `onError`, for session -errors. `sender.metrics` returns a snapshot of the sender's counters, including -`metrics.ingress` with the replay queue, reconnect, and notification counters. - -## Configuration reference - -The [connect string reference](/docs/connect/clients/connect-string/) documents -every key. The Node.js client's defaults and deviations: - -| Key | Default | Notes | -|---|---|---| -| `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | -| `username`, `password`, `token` | none | Basic or bearer authentication. | -| `tls_verify`, `tls_roots` | `on`, Node.js CA bundle | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | -| `connect_timeout`, `auth_timeout_ms` | `15000` | DNS and TCP/TLS connection, and upgrade deadlines, in milliseconds. See [Connection timeouts](#connection-timeouts). | -| `auto_flush` | `on` | Master switch for the three triggers. | -| `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | -| `auto_flush_interval` | `100` | Milliseconds. `0` disables. `off` is rejected. | -| `auto_flush_bytes` | disabled | Size, or `off`. | -| `close_flush_timeout_millis` | `5000` | ACK wait in a standalone sender's `close()`. | -| `transaction` | `off` | Keep auto-flushed batches in an open transaction until `flush()`. | -| `request_durable_ack` | `off` | Enterprise. | -| `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | -| `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | -| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in default memory mode. `0` removes it. | -| `max_frame_rejections`, `poison_min_escalation_window_millis` | `4`, `300000` | Poison-frame detector: rejections of one batch, and the minimum time they must span, before the sender stops. | -| `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | -| `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | -| `sf_durability` | `memory` | `memory`, `periodic`, or `append`. | -| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](#sf-capacity), not a hard disk limit; memory queue cap without `sf_dir`. | -| `sf_max_segment_bytes` | `4m` with `sf_dir`, none without | Journal segment size, which also caps a batch. | -| `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | -| `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. | -| `target`, `zone` | `any`, none | Endpoint role and zone preference. Apply to ingestion too. | -| `failover`, `failover_max_attempts`, `failover_max_duration_ms` | `on`, `8`, `30000` | Query failover. | -| `compression`, `compression_level` | `raw`, `1` | Query result compression. | -| `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | -| `client_id` | `typescript/` | Sent to the server for diagnostics. | -| `error_inbox_capacity`, `connection_listener_inbox_capacity` | `256`, `64` | Queues for rejection callbacks and connection events. | -| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client. A standalone `Sender` also applies `lazy_connect`. | - -The -[API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) -covers every type and option. The -[QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) -in the client repository describes the delivery semantics in depth. - -### Differences from other clients - -The Node.js client differs from the Java reference client, and from the shared -[connect string reference](/docs/connect/clients/connect-string/), in these -places: - -| Area | Node.js behavior | -|---|---| -| Outage budget | A sender in default memory mode gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. Senders in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) or with `sf_dir` retry indefinitely. See [Ingestion reconnect](#ingestion-reconnect). | -| `target` and `zone` | Also apply to ingestion. Set a query-only role with the typed `egress.target` option. See [Multiple endpoints](#multiple-endpoints). | -| Authentication rejected after a first connection | Senders with `sf_dir` or in background memory mode keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | -| Durable acknowledgement unavailable | Senders in background memory mode keep retrying from startup, and store-and-forward senders after their first connection, emitting `durable-ack-unavailable`. See [Durable acknowledgement](#durable-acknowledgement). | -| `sf_durability` | Also accepts `append`. | -| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Journal capacity](#sf-capacity). | -| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Lock recovery](#sf-lock-recovery). | -| `max_lifetime_ms` | Closes idle connections above the pool minimum only. Connections at the minimum are not recycled. | -| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush or from sender creation. Size values take single-letter suffixes only. `tls_roots` must be PEM. `init_buf_size` and `max_buf_size` are rejected. | -| `connect_timeout` | Also covers DNS and the TLS handshake, and `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. See [Connection timeouts](#connection-timeouts). | -| `tls_roots` default | The CA certificates bundled with Node.js, not the operating system's trust store. See [TLS](#tls). | -| Defaults | `connect_timeout` is `15000` and `poison_min_escalation_window_millis` is `300000`. `close_flush_timeout_millis` is `5000`, as in the Rust, C, C++, Python, and Go clients; Java and .NET use `60000`. | -| Close after an ACK timeout | A standalone sender's `close()` rejects with `QwpSenderCloseTimeoutError` instead of logging a warning. The pooled client's `db.close()` resolves and reports the timeout, best-effort, to `ingressSession.onError`. See [Closing a sender](#closing-a-sender). | -| Error reports | Categories and policies are lowercase, hyphenated strings, such as `schema-mismatch` and `retriable-other`. See [Ingestion errors](#ingestion-errors). | -| Pool and query keys on a standalone `Sender` | The `Sender` logs a warning for the pool and query-only keys it ignores. It applies `client_id` and `lazy_connect`. | -| `on_*_error` keys | Accepted but not applied. | - -## Migration - -### From ILP to QWP - -The row API is unchanged, so existing `Sender` code migrates by changing the -connect string and calling `connect()`: - -```diff -- const sender = await Sender.fromConfig("http::addr=localhost:9000"); -+ const sender = await Sender.fromConfig("ws::addr=localhost:9000"); -+ await sender.connect(); -``` - -| Aspect | ILP over HTTP | QWP over WebSocket | -|---|---|---| -| Connect string schema | `http::`, `https::` | `ws::`, `wss::` | -| Auto-flush rows | 75,000 (600 over TCP) | 1,000 | -| Auto-flush interval | 1,000 ms | 100 ms | -| `flush()` completes when | QuestDB responds to the HTTP request | The batch is published; the ACK arrives later | -| Server rejection | `flush()` throws | Asynchronous: `onSenderError`, `waitForAcknowledged()`, or `flush()` with `awaitServerAck` | -| Rows staged at `close()` | Lost unless flushed | Published; waits up to 5 seconds for ACK, then unacknowledged rows may be lost without `sf_dir` | -| Reconnect and replay | Retries one request for `retry_timeout` | Automatic, with replay of unacknowledged batches | -| Store-and-forward, querying, pooling | Not available | Available | -| Column types | ILP types | More types, subject to [column-method](#column-methods) and [array](#arrays) support | - -Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, -`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss` -with a hint. It names the replacement where there is one -(`retry_timeout` becomes `reconnect_max_duration_millis`, and `tls_ca` becomes -`tls_roots`), and otherwise says that the key applies only to ILP or that QWP -negotiates the setting itself. To keep ILP-sized batches, set -`auto_flush_rows` and `auto_flush_interval` explicitly. Migrate one sender at a -time: ILP and QWP senders can run side by side. - -### Upgrading from 4.x - -Version 5.0.0 keeps the ILP API and adds QWP. Changes that affect existing ILP -code: - -- **Null values.** Passing `null` or `undefined` to a column or symbol method now - omits the column. Existing nullable columns store NULL; BOOLEAN defaults to - `false`, and BYTE and SHORT default to `0` (see [Null values](#null-values)). - Earlier versions threw a type error for most such values. Validate data - before calling the sender if you relied on the error. -- **Decimal scale.** `decimalColumn()` over ILP rejects a non-integer `scale` - with a `RangeError`. Earlier versions silently coerced it, writing `2.5` as - scale 2 and `NaN` as scale 0. -- **`intColumn()`** also accepts a `bigint`, for LONG values beyond - `Number.MAX_SAFE_INTEGER`. -- **TCP authentication** now works on Node.js 26, which rejects the JWK the - client previously built. -- **New dependency.** The package now depends on `ws`, used for QWP. - -## ILP transports (legacy) - -The Node.js `Sender` still ingests over ILP, for existing deployments and for -servers without QWP. To move ILP code to QWP, see -[From ILP to QWP](#from-ilp-to-qwp); for behavior changes in 5.0.0, see -[Upgrading from 4.x](#upgrading-from-4x). ILP senders support HTTP (`http::`, -`https::`) and TCP (`tcp::`, `tcps::`) transports: - -```typescript -import { Sender } from "@questdb/nodejs-client"; - -const sender = await Sender.fromConfig( - "http::addr=localhost:9000;username=admin;password=quest;", -); -try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.00044) - .at(Date.now(), "ms"); - // ILP does not flush on close: rows still buffered at close() are lost. - await sender.flush(); -} finally { - await sender.close(); -} -``` - -- HTTP connects per request, so `connect()` is not needed; TCP transports - require `await sender.connect()`. `token=...` selects bearer authentication - over HTTP. Over TCP, `username` and `token` set the JWK key ID and private - key. -- Over HTTP, `flush()` sends the buffer as one request and throws if QuestDB - rejects it. Data is transactional only for a single-table request. A - multi-table request can commit earlier tables before a later table fails, - so a failed flush does not mean no data was committed. Schema changes, such - as automatically added columns, are not rolled back even for a single-table - request. See [HTTP transaction semantics](/docs/connect/compatibility/ilp/overview/#http-transaction-semantics). -- Decimals need ILP protocol version 3: HTTP negotiates it automatically, and - TCP needs `protocol_version=3`. Arrays need version 2 or later. -- Undici is the default HTTP agent. Set `stdlib_http=on` to use the Node.js - `http` module instead. - -For ILP options, see the -[`SenderOptions` reference](https://questdb.github.io/nodejs-questdb-client/classes/_questdb_nodejs-client.SenderOptions.html) -and the [ILP overview](/docs/connect/compatibility/ilp/overview/). - -## Full example: Ingestion and querying with failover - -A production-oriented pattern that ingests trades and queries recent prices, -with TLS, a token, several hosts, error handling, and failover handling. Before -running it, create the deduplicated table on the primary (or reuse the table -from [Store-and-forward](#store-and-forward)). If the table is missing, QWP -creates it without deduplication: - -```questdb-sql -CREATE TABLE IF NOT EXISTS trades_sf ( - timestamp TIMESTAMP, - trade_id VARCHAR, - symbol SYMBOL, - side SYMBOL, - price DOUBLE, - amount DOUBLE -) TIMESTAMP(timestamp) PARTITION BY DAY -DEDUP UPSERT KEYS(timestamp, trade_id); -``` - -Replace the sample events with source-assigned trade IDs and timestamps. Keep -both values unchanged when retrying the same event, and use a writable, -persistent `sf_dir` so unacknowledged rows survive a shutdown: - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, - QwpIngressAckTimeoutError, - QwpPoolResourceError, - QWP_RECONNECT_EVENT_KIND, - QWP_SENDER_ERROR_POLICY, - type QwpReconnectEvent, - type QwpSenderError, -} from "@questdb/nodejs-client"; - -const token = process.env.QDB_TOKEN; -if (!token) throw new Error("QDB_TOKEN is not set"); - -// This example runs one query at a time so its replay callback can clear it. -const recentPrices: (readonly unknown[])[] = []; - -// Replace with your alerting. -function alertOperator(message: string) { - console.error("ALERT:", message); -} - -function logConnection(event: QwpReconnectEvent) { - if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { - const endpoint = String(event.endpoint ?? ""); - console.info("questdb connection:", event.kind, endpoint); - } -} - -const db = await connectQwpNodeClient( - "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + - `token=${token};` + - // append: every flush waits for a disk sync; see "Store-and-forward". - "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + - // Limit offline batches below the default server's 2 MiB limit. - "sf_durability=append;sf_max_segment_bytes=1m;sender_pool_max=4;" + - // Query pool stays cold while replicas are down; one active query so the - // replay callback below can reset its state even if replay has no batches. - "query_pool_min=0;query_pool_max=1;", - { - // Queries run on replicas only, never on the primary; ingestion always - // follows the primary. - egress: { target: "replica", compression: "zstd" }, - ingressSession: { - onSenderError: (error: QwpSenderError) => { - console.error("batch rejected:", error.category, error.serverMessage); - // A terminally rejected batch stays in the journal and blocks - // ingestion through this client, for every table, until it is fixed. - if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { - alertOperator(`QuestDB rejected a batch: ${error.serverMessage}`); - } - }, - // Terminal failures, such as a batch that QuestDB rejects terminally. - onError: (event) => { - if (event.terminal) console.error("ingestion stopped:", event.error); - }, - // Replaces any reconnect_* keys; omitted fields use the defaults. - reconnect: { onEvent: logConnection }, - }, - egressSession: { - queryTimeoutMs: 30_000, - // Replaces any failover* keys; omitted fields use the defaults. - // No 8-attempt limit (maxAttempts 0): failover can last up to 30 s. - reconnect: { - maxAttempts: 0, - maxDurationMs: 30_000, - onEvent: logConnection, - }, - onReplayReset: (event) => { - recentPrices.length = 0; - console.warn("query restarts on", String(event.endpoint)); - }, - }, - }, -); - -try { - // Ingestion: one borrowed sender per producer. IDs and timestamps must - // come from the source, not be regenerated on an application retry. - const events = [ - { - tradeId: "trade-12345", - timestampMs: 1723000000000, - symbol: "ETH-USD", - price: 2615.54, - amount: 0.5, - }, - { - tradeId: "trade-12346", - timestampMs: 1723000000001, - symbol: "BTC-USD", - price: 39269.98, - amount: 0.001, - }, - ]; - const sender = await db.borrowSender(); - try { - for (const event of events) { - await sender - .table("trades_sf") - .stringColumn("trade_id", event.tradeId) - .symbol("symbol", event.symbol) - .symbol("side", "buy") - .doubleColumn("price", event.price) - .doubleColumn("amount", event.amount) - .at(event.timestampMs, "ms"); - } - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence, 10_000); - } catch (error) { - if (!(error instanceof QwpIngressAckTimeoutError)) throw error; - console.warn("ACK timed out; rows remain in sf_dir for replay after close"); - } finally { - // After a terminal rejection, close() rejects with the same failure. - // Log it so that it does not replace the error thrown above. - await sender - .close() - .catch((error) => console.error("close failed:", error)); - } - - // Querying: rows may not be visible yet, see "Read-after-write". - try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query( - "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + - "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", - { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, - ); - for await (const batch of query) { - // A nonempty replay starts at sequence 0; the callback handles - // empty ones. - if (batch.batchSequence === 0n) recentPrices.length = 0; - for (const row of batch.rows()) recentPrices.push(row); - } - await query.completion; - console.log(recentPrices); - } finally { - await lease.close(); - } - } catch (error) { - if (error instanceof QwpPoolResourceError) { - // No replica was reachable within the failover budget. - console.warn("no replica available for queries:", error.cause); - } else if (error instanceof QwpEgressQueryError) { - console.error(`query failed: status=${error.status} ${error.message}`); - } else { - throw error; - } - } -} finally { - await db.close(); -} -``` - -The query can still miss newly acknowledged rows until WAL apply catches up; -use the [Read-after-write](#read-after-write) pattern for a visibility guarantee. -A replayed batch is idempotent only because this example retains the event's -ID and timestamp and enables table-level deduplication. +For typed options and the key table, see the +[Node.js configuration reference](/docs/connect/clients/nodejs-operations/#configuration-reference). ## Next steps +- [Node.js operations and reference](/docs/connect/clients/nodejs-operations/) + for pool sizing, failover, error handling, configuration, migration, and a + complete ingestion and querying example. + - [Connect string reference](/docs/connect/clients/connect-string/) for every configuration key. - [Delivery semantics](/docs/concepts/delivery-semantics/) for at-least-once diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index c5ae8c9fae..d924d3260d 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -141,7 +141,7 @@ retryable failures and stay within the failover budget. With no explicit policy, failover is enabled for established query connections but initial connection attempts are not retried. `failover=off` disables the reconnect wrapper; an explicit `egressSession.reconnect` value overrides the connect-string policy. -See [Node.js connection events](/docs/connect/clients/nodejs/#connection-events). +See [Node.js connection events](/docs/connect/clients/nodejs-operations/#connection-events). ::: @@ -374,11 +374,12 @@ The Node.js client retries authentication rejections after the first connection only in senders with `sf_dir` or in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`); see [Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). -For unsupported durable acknowledgement, its senders in background memory -mode keep retrying from startup -and emit `durable-ack-unavailable`, and store-and-forward senders do the same -after their first successful connection. Monitor these -[connection events](/docs/connect/clients/nodejs/#connection-events) and buffer +For unsupported durable acknowledgement, Node.js senders with a background +start (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying from +startup and emit `durable-ack-unavailable`, even with `sf_dir`. With a +foreground start and `sf_dir`, the first connection fails but later mismatches +are retried after a successful connection. Monitor these +[connection events](/docs/connect/clients/nodejs-operations/#connection-events) and buffer usage; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). ### Reconnect and outage handling @@ -410,8 +411,8 @@ This is the behaviour of the Java reference client and the .NET client. Other clients are aligned to it, except a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, which gives up after `reconnect_max_duration_millis`; see the -[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect) and its -[other differences](/docs/connect/clients/nodejs/#differences-from-other-clients). +[Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect) and its +[other differences](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). If you are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, @@ -429,7 +430,7 @@ and apply back-pressure to the producer rather than dropping data. | HTTP upgrade timeout / non-auth transport error | try next endpoint | | `421` with `X-QuestDB-Role: REPLICA` | role reject; try next endpoint | | `401` / `403` auth failure | never try later endpoints; **terminal** before the first successful connection, then client-specific: [Java and some Node.js senders retry](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) ⚠ | -| durable-ack requested but unsupported | terminal mismatch, except that Java senders retry after their first successful connection, and Node.js senders retry from startup in background memory mode, or after their first connection with `sf_dir` ([details](/docs/connect/clients/nodejs/#durable-acknowledgement)) | +| durable-ack requested but unsupported | terminal mismatch, except that Java senders retry after their first successful connection. Node.js senders retry from startup with a background start (with or without `sf_dir`); with foreground startup and `sf_dir`, they retry only after a first successful connection ([details](/docs/connect/clients/nodejs/#durable-acknowledgement)) | | successful write upgrade | bind this endpoint | | all endpoints fail transport | throw / retry per initial/reconnect mode | | all endpoints role-reject as replicas | `QwpRoleMismatchException` | diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index 297a2cb2e5..18321f7bc5 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1151,7 +1151,7 @@ section of the connect string reference: | Key | Default | Description | |----------------------------------|-----------|-------------------------------------------| -| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on` ([details](/docs/connect/clients/nodejs/#differences-from-other-clients)). | +| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on` ([details](/docs/connect/clients/nodejs-operations/#differences-from-other-clients)). | | `reconnect_initial_backoff_millis` | `100` | First post-failure sleep. | | `reconnect_max_backoff_millis` | `5000` | Cap on per-attempt sleep. | | `initial_connect_retry` | `off` | Retry on first connect (`on`, `sync`, `async`). | diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index 0a2bb2db11..db808f3347 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -83,7 +83,7 @@ both storage modes, so the `zone=` key is silently accepted on ingress connections and only takes effect on egress. The Node.js client is the exception: it applies `zone=` and `target=` to ingress too. Its other deviations are listed under -[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). ### Selection priority @@ -155,7 +155,7 @@ length, and what bounds your tolerance is buffer capacity `initial_connect_retry=async`, or `lazy_connect=on`, which gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`; see the - [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). + [Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps reconnect storms when many producers share a cluster - Inter-host pause within a round: **none** — the client walks the full @@ -264,7 +264,7 @@ its buffer capacity, so a credential change on the cluster does not stop the producer. Monitor such a sender: the Java client reports each rejection to the sender's error handler as a retriable `SECURITY_ERROR`, and the Node.js client emits an `attempt-failed` connection event for each failed attempt. See -[Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors) +[Node.js connection errors](/docs/connect/clients/nodejs-operations/#connection-level-errors) for the Node.js rules. Per-host credentials are outside the failover model. Use a separate connect diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 0d751ea8c0..dda81c40cd 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -19,7 +19,7 @@ first. ingress parsers accept and ignore them, so one connect string can serve both directions. The Node.js client is the exception: it applies both keys to ingress too; its other deviations are listed under -[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). They are documented in full on the [connect-string reference](/docs/connect/clients/connect-string#failover-keys); the table below summarises the failover-relevant subset. @@ -27,8 +27,8 @@ the table below summarises the failover-relevant subset. | Key | Type | Default | Notes | |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | -| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only** (the Node.js client also applies it to ingress). Which server role the query client accepts. Other clients accept and ignore it on an ingress connect string. On the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), set the query-side role through the typed `egress` option instead. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs-operations/#multiple-endpoints), which also ranks ingress endpoints by zone. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only** (the Node.js client also applies it to ingress). Which server role the query client accepts. Other clients accept and ignore it on an ingress connect string. On the [Node.js client](/docs/connect/clients/nodejs-operations/#multiple-endpoints), set the query-side role through the typed `egress` option instead. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | | `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. `connect_timeout` bounds the TCP connect separately: most clients leave it unset by default and then use the OS timeout, while Node.js defaults it to 15 s and also bounds DNS and TLS with it. On Node.js, `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. Lower `auth_timeout_ms` for faster upgrade failure detection; tune the connect timeout separately. | `addr` syntax — both of these are equivalent and produce the same three-peer @@ -50,7 +50,7 @@ for the full list. The failover-relevant keys are: | Key | Type | Default | Notes | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index 7a3ea586a8..e162e74035 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -45,8 +45,8 @@ memory mode in two. A sender in default memory mode gives up after `reconnect_max_duration_millis` (5 minutes by default), while a sender in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) retries indefinitely, as SF mode does. See the -[Node.js ingestion modes](/docs/connect/clients/nodejs/#flushing) and -[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). +[Node.js ingestion modes](/docs/connect/clients/nodejs/#ingestion-modes) and +[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). ## What "frame" means here @@ -140,10 +140,11 @@ object store** (S3, Azure Blob, GCS, or NFS). frames it cannot receive. In most clients the rejection is terminal. The Java client retries it after a sender's first successful connection, so a capability change on the cluster cannot stop the producer. Node.js senders - in background memory mode (`initial_connect_retry=async` or - `lazy_connect=on`) retry it from startup, and Node.js store-and-forward - senders after their first - connection, emitting `durable-ack-unavailable` events; see + with a background start (`initial_connect_retry=async` or + `lazy_connect=on`) retry from startup and emit `durable-ack-unavailable`, + with or without `sf_dir`. A Node.js sender with `sf_dir` and a foreground + start fails on the first connection but retries after a successful + connection; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). A retrying sender keeps buffering, so monitor it. @@ -168,7 +169,7 @@ On Node.js, only a sender in default memory mode waits for the reconnect in `flush()`, up to `reconnect_max_duration_millis`. Background memory mode, enabled by `initial_connect_retry=async` or `lazy_connect=on`, keeps accepting batches into the memory replay queue until capacity is exhausted and retries -indefinitely. See the [three Node.js ingestion modes](/docs/connect/clients/nodejs/#flushing). +indefinitely. See the [three Node.js ingestion modes](/docs/connect/clients/nodejs/#ingestion-modes). On every successful (re)connect: diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 4a1c3158cf..345e707e52 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -48,7 +48,7 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 7d862e7f40..3bd147a055 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -60,7 +60,7 @@ switch between them without changing application code — only the connect string. On the Node.js client, a sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, also gives up after `reconnect_max_duration_millis` of outage; see -[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). ## Comparison at a glance @@ -114,12 +114,12 @@ GCS, or NFS). or an uninitialised primary, the connection attempt is rejected. In most clients this is terminal. The exceptions follow. - **Senders that keep retrying.** The Java client retries after a sender's - first successful connection. Node.js senders in background memory mode - (`initial_connect_retry=async` or `lazy_connect=on`) retry from startup - instead of failing initialization, and Node.js store-and-forward senders - retry after their first successful connection; they emit - `durable-ack-unavailable` - [connection events](/docs/connect/clients/nodejs/#connection-events). + first successful connection. Node.js senders with a background start + (`initial_connect_retry=async` or `lazy_connect=on`) retry from startup, + even with `sf_dir`. With `sf_dir` and a foreground start, the first + connection fails but later mismatches are retried after a successful + connection. Retrying Node.js senders emit `durable-ack-unavailable` + [connection events](/docs/connect/clients/nodejs-operations/#connection-events). Monitor retrying senders and their buffer usage: continued buffering can fill the journal or memory queue even though startup succeeded. See [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). diff --git a/documentation/query/overview.md b/documentation/query/overview.md index 9167063fb5..71317d0433 100644 --- a/documentation/query/overview.md +++ b/documentation/query/overview.md @@ -119,9 +119,9 @@ SQL writes such as `INSERT`. Each client describes its failover behavior: [Python](/docs/connect/clients/python/#reader-failover), [Go](/docs/connect/clients/go/#query-failover), [Rust](/docs/connect/clients/rust/#failover-and-errors), -[C and C++](/docs/connect/clients/c-and-cpp/#failover-retry-and-pool-lifecycle), -[.NET](/docs/connect/clients/dotnet/#failover-and-high-availability), and -[Node.js](/docs/connect/clients/nodejs/#query-failover). +[C and C++](/docs/connect/clients/c-and-cpp/#querying-data), +[.NET](/docs/connect/clients/dotnet/#failover), and +[Node.js](/docs/connect/clients/nodejs-operations/#query-failover). The Rust, C++, and Python clients hand back results as Arrow record batches. That is the native memory layout of diff --git a/documentation/sidebars.js b/documentation/sidebars.js index e824810980..1b91bfb023 100644 --- a/documentation/sidebars.js +++ b/documentation/sidebars.js @@ -78,6 +78,11 @@ module.exports = { type: "doc", label: "Node.js", }, + { + id: "connect/clients/nodejs-operations", + type: "doc", + label: "Node.js operations and reference", + }, { id: "connect/clients/c-and-cpp", type: "doc", diff --git a/shared/clients.json b/shared/clients.json index 393030b9df..86d54ac4d8 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -112,7 +112,7 @@ "protocol": "PGWire" }, { - "href": "/docs/connect/clients/nodejs#ilp-transports-legacy", + "href": "/docs/connect/clients/nodejs-operations#ilp-transports-legacy", "name": "Node.js", "description": "Node.js client for ILP ingestion over HTTP and TCP, alongside QWP.", From db92c89f63b4149192515a9d0bdbe222c0eefba1 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Thu, 1 Oct 2026 17:44:04 +0100 Subject: [PATCH 18/25] docs: simplify Node.js client guide into one page --- documentation/concepts/delivery-semantics.md | 2 +- .../connect/clients/connect-string.md | 8 +- .../connect/clients/nodejs-operations.md | 1413 --------- documentation/connect/clients/nodejs.md | 2636 ++++------------- .../wire-protocols/qwp-client-behavior.md | 8 +- .../wire-protocols/qwp-ingress-websocket.md | 2 +- .../client-failover/concepts.md | 6 +- .../client-failover/configuration.md | 8 +- .../store-and-forward/concepts.md | 2 +- .../store-and-forward/configuration.md | 2 +- .../store-and-forward/when-to-use.md | 4 +- documentation/query/overview.md | 2 +- documentation/sidebars.js | 5 - .../raw-markdown/convert-components.test.js | 2 +- shared/clients.json | 2 +- 15 files changed, 595 insertions(+), 3507 deletions(-) delete mode 100644 documentation/connect/clients/nodejs-operations.md diff --git a/documentation/concepts/delivery-semantics.md b/documentation/concepts/delivery-semantics.md index ec6c3dc439..201651a5a4 100644 --- a/documentation/concepts/delivery-semantics.md +++ b/documentation/concepts/delivery-semantics.md @@ -19,7 +19,7 @@ unacknowledged rows live in memory and are lost if the process exits, or the sender closes, before the server acknowledges them. A Node.js sender in default memory mode also gives up after `reconnect_max_duration_millis` of outage; see the -[Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). +[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). This page explains where duplicates come from and how to suppress them. diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index c92fdbce79..9b83453ad8 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -20,7 +20,7 @@ configures both without edits. The Node.js client is the exception for `target` and `zone`, which it also applies to ingress; see [Role filter and zone preference](#role-filter-and-zone-preference). The *Applies to:* tag on each section below marks which direction a key affects. -The [Node.js client page](/docs/connect/clients/nodejs-operations/#differences-from-other-clients) +The [Node.js client page](/docs/connect/clients/nodejs/#differences-from-other-clients) lists its behavioral differences from this reference. For legacy InfluxDB Line Protocol (ILP) transports (`http`, `https`, `tcp`, @@ -454,7 +454,7 @@ The Node.js client applies `target` and `zone` to ingress as well. With `target=replica` in a shared connect string, its senders accept only replicas and cannot ingest. Set the query-side role through the typed `egress` option instead; see the -[Node.js client page](/docs/connect/clients/nodejs-operations/#multiple-endpoints). +[Node.js client page](/docs/connect/clients/nodejs/#multiple-endpoints). ::: @@ -672,7 +672,7 @@ exception is a Node.js sender in default memory mode, which gives up after `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies this budget to every outage, and fails with `QwpReconnectExhaustedError` when it runs out. See the - [Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). + [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - `initial_connect_retry` — whether the client retries the initial connect attempt on failure. - `off` (default, alias `false`) — fail fast on initial connect failure. @@ -757,7 +757,7 @@ transport-level OK ACK alone cannot close. events. Default: `200` (ms). Set to `0` or a negative value to disable in clients that support it. In Node.js, explicitly setting this key also requests durable ACK (even at `0`), and negative values are rejected; - see [Node.js differences](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). + see [Node.js differences](/docs/connect/clients/nodejs/#differences-from-other-clients). See the [QWP Egress (WebSocket)](/docs/connect/wire-protocols/qwp-egress-websocket/) wire protocol for the underlying mechanism. diff --git a/documentation/connect/clients/nodejs-operations.md b/documentation/connect/clients/nodejs-operations.md deleted file mode 100644 index 2e5a40b8a6..0000000000 --- a/documentation/connect/clients/nodejs-operations.md +++ /dev/null @@ -1,1413 +0,0 @@ ---- -slug: /connect/clients/nodejs-operations -title: Node.js client operations and reference -sidebar_label: Node.js operations and reference -description: "Node.js QWP client pool lifecycle, error recovery, failover, configuration, migration, and a complete ingestion and query example." ---- - -For a first connection, row ingestion, and streaming SQL queries, start with the -[Node.js client guide](/docs/connect/clients/nodejs/). This companion page -covers pool lifecycle, concurrency, error recovery, failover, connect-string -differences, migration, and a complete ingestion and query example. - -## The connection pool - -The pooled client keeps two elastic pools: one of senders and one of query -connections. Each pool opens its minimum on `connect()`, grows on demand up to -its maximum, and a housekeeper closes connections that stay idle too long or -exceed their maximum lifetime, never going below the minimum. - -### Borrowing a sender - -A borrowed sender belongs to the borrower until its `close()` returns it: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - for (const [symbol, price] of [ - ["ETH-USD", 2615.54], - ["BTC-USD", 39269.98], - ] as const) { - await sender - .table("trades") - .symbol("symbol", symbol) - .symbol("side", "buy") - .doubleColumn("price", price) - .doubleColumn("amount", 0.1) - .at(Date.now(), "ms"); - } - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence); - } finally { - // Returns the sender to the pool; any pending rows are flushed. - await sender.close(); - } -} finally { - await db.close(); -} -``` - -A long-running producer can keep its borrow for its whole lifetime and call -`flush()` between batches. Size `sender_pool_max` to the number of producers -that hold a sender at the same time. - -`close()` on a borrowed sender flushes its completed rows and returns it to -the pool. The example above waits for the acknowledgement *before* returning -the sender: `close()` itself does not wait for acknowledgements, although in -default memory mode its flush waits for the reconnect during an outage. See -[Closing a borrowed sender](/docs/connect/clients/nodejs/#closing-a-borrowed-sender) for how long that can -take, how to wait for acknowledgements, and what happens when a close fails. - -After `close()`, every method call or property read on that sender object -throws `QwpClientClosedError`. Don't keep references to a returned sender, for -example in callbacks that can run later. - -### Borrowing a query lease - -A query lease runs one query at a time. For concurrent queries, borrow one lease -per query, up to `query_pool_max`: - -```typescript -import { - connectQwpNodeClient, - type QwpQueryLease, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); - -async function countBySymbol(lease: QwpQueryLease, symbol: string) { - const query = await lease.query( - "SELECT count() FROM trades WHERE symbol = $1", - { binds: (binds) => binds.setVarchar(0, symbol) }, - ); - let count = 0n; - for await (const batch of query) count = batch.get(0, 0) as bigint; - await query.completion; - return count; -} - -try { - const [a, b] = await Promise.all([db.borrowQuery(), db.borrowQuery()]); - try { - // Two leases, two WebSockets: the queries run concurrently. - const [eth, btc] = await Promise.all([ - countBySymbol(a, "ETH-USD"), - countBySymbol(b, "BTC-USD"), - ]); - console.log({ eth, btc }); - } finally { - await Promise.all([a.close(), b.close()]); - } -} finally { - await db.close(); -} -``` - -Starting a second query on a lease while one is still active rejects with -`a QWP query is already active on this connection`. Always close a lease in -`finally`: an unreturned lease holds its connection until `db.close()`. - -### Pool settings - -| Key | Default | Purpose | -|---|---|---| -| `sender_pool_min` | `1` | Senders kept open even when idle. `0` lets the pool close them all. | -| `sender_pool_max` | `4` | Maximum senders the pool opens. | -| `query_pool_min` | `1` | Query connections kept open even when idle. | -| `query_pool_max` | `4` | Maximum query connections, which also caps concurrent queries. | -| `acquire_timeout_ms` | `5000` | How long a borrow waits when the pool is at its maximum, before rejecting with `QwpPoolAcquireTimeoutError`. | -| `idle_timeout_ms` | `60000` | Idle time before an excess connection is closed. `0` keeps idle connections. | -| `max_lifetime_ms` | `1800000` | Age at which an idle connection above the pool minimum is closed. Connections kept open by `sender_pool_min` and `query_pool_min` are never recycled, so this does not rotate a pool that is at its minimum. `0` disables it. | -| `housekeeper_interval_ms` | `5000` | How often the housekeeper checks for idle and over-age connections. Minimum `100`. | -| `query_close_timeout_ms` | `5000` | How long returning a lease with an active query waits for the cancellation to drain before discarding the connection. | -| `lazy_connect` | `off` | Start without connecting. See below. | - -Pool sizes, acquisition and idle timeouts, lifetime, and housekeeping settings -have typed equivalents in the `pool` section of the second argument -(`senderPoolMin`, `acquireTimeoutMs`, `housekeepingIntervalMs`, and so on). -The other two settings use different locations: - -- `query_close_timeout_ms` maps to `egressSession.cancelDrainTimeoutMs`, not - `pool`. -- Set `lazy_connect=on` in the connect string. When passing a full - `QwpNodeClientOptions` object instead of a string, use top-level - `lazyConnect: true`. It is not supported in `pool` or the second argument. - -When creating a new pooled connection fails, the borrow rejects with -`QwpPoolResourceError`, whose `cause` holds the connection error. - -`borrowSender()` and `borrowQuery()` take no timeout argument. When the pool -is at its maximum, a borrow waits up to `acquire_timeout_ms` for a connection -to be returned. Opening a new connection is bounded by the -[connection timeouts](#connection-timeouts) of each endpoint, and by the -[failover budget](#query-failover) when query retries are on. To enforce a -shorter deadline, such as a request deadline, race the borrow against a timer -and return a lease that arrives late: - -```typescript -import { connectQwpNodeClient, type QwpClient } from "@questdb/nodejs-client"; - -function borrowQueryWithin(db: QwpClient, timeoutMs: number) { - const borrow = db.borrowQuery(); - let timer: ReturnType | undefined; - const deadline = new Promise((_, reject) => { - timer = setTimeout(() => reject(new Error("borrow timed out")), timeoutMs); - }); - return Promise.race([borrow, deadline]) - .catch((error: unknown) => { - // Return a lease that arrives after the deadline. - borrow.then((lease) => lease.close(), () => undefined); - throw error; - }) - .finally(() => clearTimeout(timer)); -} - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await borrowQueryWithin(db, 2_000); - try { - const query = await lease.query("SELECT count() FROM trades"); - for await (const batch of query) console.log(batch.get(0, 0)); - await query.completion; - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -### Starting while QuestDB is down - -`connectQwpNodeClient()` fails fast when QuestDB is unreachable. Set -`lazy_connect=on` to start regardless: senders connect in the background and -buffer rows in memory until QuestDB is reachable. The query pool stays empty -until the first query. - -```typescript -import { - connectQwpNodeClient, - QwpIngressAckTimeoutError, -} from "@questdb/nodejs-client"; - -// Resolves immediately, even if QuestDB is not running yet. -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;lazy_connect=on;", -); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); - const sequence = sender.publishedSequence; - // Keep the process running until QuestDB comes back and acknowledges it. - for (;;) { - try { - await sender.waitForAcknowledged(sequence, 10_000); - break; - } catch (error) { - if (!(error instanceof QwpIngressAckTimeoutError)) throw error; - console.info("still waiting for QuestDB; do not restage the row"); - } - } - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -Rows buffered while QuestDB is down exist only in memory. This example keeps -the client running until the row is acknowledged; if the process exits first, -the unacknowledged row may be lost. See -[Closing the pooled client](#closing-the-pooled-client). -To keep them across a shutdown or restart, add a -[store-and-forward](/docs/connect/clients/nodejs/#store-and-forward) journal with `sf_dir`. `lazy_connect=on` -still starts the sender in the background when `sf_dir` is set; the journal -changes where rows are buffered, not whether startup waits for a connection. -Replay from the journal is at least once, so write to a deduplicated table as -described there. - -`lazy_connect=on` forces `query_pool_min=0` and `initial_connect_retry=async`, -and rejects an explicit conflicting value. Setting `initial_connect_retry=async` -without `lazy_connect` is not enough: the query pool still connects at startup, -so `connectQwpNodeClient()` rejects with `QwpPoolResourceError`. A query -borrowed while QuestDB is still down rejects with `QwpPoolResourceError` too. - -A lazy start does not cover two cases. A locked store-and-forward journal -still fails startup; see [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery). And until a -sender has connected once, it cannot check batches against the server's size -limit; see [Batch size limits](/docs/connect/clients/nodejs/#batch-size-limits). - -### Closing the pooled client - -`db.close()` rejects new borrows, then: - -- Cancels active queries and closes every query connection, including leased - ones. -- Closes idle senders. Each publishes its remaining rows and waits up to - `close_flush_timeout_millis` (5 seconds) for QuestDB to acknowledge them. -- Waits for borrowed senders to be returned, until 5 seconds after - `db.close()` was called, or `acquire_timeout_ms` if that is lower. Closing - the idle senders counts toward the same deadline. A sender still borrowed - after that stays open: its owner must `close()` it, and the process stays - alive until then. - -`db.close()` resolves even when an acknowledgement does not arrive in time. -Without `sf_dir`, the unacknowledged rows are then lost. With `sf_dir`, they -stay in the journal, and the next sender on the same directory replays them. -The client usually reports the timeout to `ingressSession.onError` as a -non-terminal `QwpIngressAckTimeoutError`, logged as a warning by default, but -the report is best-effort: do not rely on it to detect unacknowledged rows. To -know that QuestDB accepted every row before shutting down, wait for the -acknowledgement before returning each sender (see -[Awaiting acknowledgements](/docs/connect/clients/nodejs/#awaiting-acknowledgements)), or use -[store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). - -## Concurrency - -Node.js runs your code on one thread, but async functions interleave at every -`await`: - -- **`QwpClient`** is safe to share across your whole application. -- **Senders** are not safe for concurrent producers. A row is built across - several calls, so an `await` between `table()` and `at()` lets another task - add columns to the same row. Give each producer its own sender, borrowed from - the pool, and size `sender_pool_max` to match. -- **Query leases** run one query at a time. Borrow one lease per concurrent - query; `query_pool_max` caps concurrent queries. -- **Worker threads** cannot share clients. Create one client per worker, and - give each worker its own `sender_id` when using store-and-forward. - -Callbacks such as `onSenderError` run on the same event loop, so move CPU-heavy -work out of them. Row encoding runs on the event loop too, so one process -ingests at most as fast as one CPU core allows. To go faster, split the stream -across worker threads or processes, each with its own client. - -## Error handling - -Each error leaves the client in a known state. The sections after this table -have the details and examples: - -| Error | Surfaces from | State afterwards | What to do | -|---|---|---|---| -| `TypeError`, `RangeError`, or `Error` from local validation | The column method or `at()` that staged the value | The row in progress is discarded; the sender stays usable | Fix the value and write the row again | -| `QwpBatchTooLargeError` | `flush()`, the `at()` whose auto-flush sends the batch, or `close()` | The batch can never be sent: the staged rows are kept, every later flush fails the same way, and `close()` discards them | Call `reset()`, then write the rows again without the oversized one; see [Batch size limits](/docs/connect/clients/nodejs/#batch-size-limits) | -| `QwpMemoryReplayAppendTimeoutError`, `QwpReplayStoreAppendTimeoutError` | `flush()`, an auto-flushing `at()`, or `close()` | The batch stays staged; the sender stays usable | Keep the sender and flush again later. Don't write the rows again, and don't close the sender while backpressure persists; see [Backpressure](/docs/connect/clients/nodejs/#backpressure) | -| Retriable server rejection | `onSenderError` | The client resends the batch; repeated rejections become terminal | Monitor; no action needed per rejection | -| Terminal server rejection | `onSenderError`, then `QwpIngressNackError` or `QwpReplayRejectedError` from later calls | The sender has failed; rows still staged on it are lost. With `sf_dir`, the batch blocks the journal for every table | Fix the data or schema, then write the lost rows on a new sender; see [Recovering from a terminal rejection](#recovering-from-a-terminal-rejection) | -| `QwpReconnectExhaustedError` on a sender | `onError` with `terminal: true`, then the next `flush()`, `at()`, or `close()` | The sender has failed; unsent rows are lost | Borrow a new sender; see [Ingestion reconnect](#ingestion-reconnect) | -| `QwpReplayStoreLockedError` | `connectQwpNodeClient()` or a borrow, as the `cause` of `QwpPoolResourceError`; `connect()` on a standalone `Sender` | The journal could not be opened | See [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery) | -| `QwpPoolResourceError` with another `cause` | `connectQwpNodeClient()`, `borrowSender()`, or `borrowQuery()` | No connection was opened | Unwrap `cause`; see [Connection-level errors](#connection-level-errors) | -| `QwpEgressQueryError` | Query iteration and `completion` | The lease stays usable | Fix the SQL or the bind values | -| `QwpEgressQueryTimeoutError`, `QwpEgressQueryAbandonedError` | Query iteration and `completion` | The lease is busy until QuestDB confirms the cancellation | Close the lease and borrow a new one | -| `QwpEgressQueryCancelTimeoutError` | Query iteration and `completion` | The connection is closed | Close the lease and borrow a new one | -| `QwpReconnectExhaustedError` on a query | Query iteration and `completion` | The lease stays failed, even after QuestDB recovers | Close the lease and borrow a new one; see [Query failover](#query-failover) | - -`QwpReplayStoreAppendTimeoutError` and the other store-and-forward journal -errors extend `QwpReplayStoreError`, so test for the specific classes before -the base class, or branch on `error.retryable`. `true` means the failure is -temporary and the sender stays usable. `false` means the journal itself can no -longer be used, for example `QwpReplayStoreCorruptionError` or -`QwpReplayStoreLockLostError`, and the sender has failed. The other classes in -the table have no common base class: test each with `instanceof`. - -### Ingestion errors - -Ingestion reports errors in two ways: - -- **While building a row.** A column method throws, or the promise returned by - `at()` or `atNow()` rejects, with a `TypeError`, `RangeError`, or `Error` for - an invalid value or name. The row in progress is discarded, and the sender - stays usable. -- **Asynchronously, when QuestDB rejects a batch.** The rejection arrives after - `flush()` resolved. It is delivered to the `onSenderError` callback, and - surfaces as a rejection of `waitForAcknowledged()`, or of `flush()` with - `awaitServerAck`. - -```typescript -import { - connectQwpNodeClient, - QWP_SENDER_ERROR_POLICY, - type QwpSenderError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - ingressSession: { - onSenderError: (error: QwpSenderError) => { - // serverStatusByte is absent for client-side errors. - const status = - error.serverStatusByte === undefined - ? "none" - : `0x${error.serverStatusByte.toString(16)}`; - console.error( - `rejected [${error.category}, policy=${error.appliedPolicy}, ` + - `status=${status}, frames=${error.fromFsn}..${error.toFsn}]: ` + - error.serverMessage, - ); - if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { - // The sender stopped: alert, and fix the data or the schema. - } - }, - }, -}); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence, 10_000); - } finally { - // Rethrows a terminal error. The pool then replaces the sender. - await sender.close(); - } -} finally { - await db.close(); -} -``` - -When `onSenderError` is not set, rejections are logged: retriable ones at -`warn`, terminal ones at `error`. Callbacks run asynchronously, never inside the -client's protocol handling, and an exception thrown by a callback is contained. -A standalone sender's `close()` can also reject, with -`QwpSenderCloseTimeoutError`, when its rows are not acknowledged in time; see -[Closing a sender](/docs/connect/clients/nodejs/#closing-a-sender). - -`QwpSenderError` fields: - -| Field | Type | Meaning | -|---|---|---| -| `category` | `string` | `schema-mismatch`, `parse-error`, `security-error`, `write-error`, `internal-error`, `not-writable`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `protocol-violation`, `data-loss`, or `unknown`. Branch on this field. | -| `appliedPolicy` | `string` | What the client did: `retriable` (reconnect and resend), `retriable-other` (resend to another endpoint), `terminal` (the sender stopped), or `abandoned` (journaled data was quarantined). | -| `serverStatusByte` | `number` | The raw QWP status code, for example `0x03` for a schema mismatch. Absent for client-side errors. | -| `serverMessage` | `string` | QuestDB's error text, for example `cannot parse DOUBLE from string [value=abc, column=price]`. | -| `fromFsn`, `toFsn` | `bigint` | The rejected frame sequence range, in the same numbering as `publishedSequence`. | -| `messageSequence` | `bigint` | The wire sequence of the rejected message. | -| `tableName` | `string` | The table, when the server attributes the rejection to one. Often absent. | -| `detectedAtMs` | `number` | When the client received the rejection. | -| `quarantinedPath` | `string` | For `data-loss` in store-and-forward: where the unreplayable journal was preserved. | - -The default policy follows the category: - -| Category | Policy | Examples | -|---|---|---| -| `schema-mismatch`, `parse-error`, `security-error`, `protocol-violation` | Terminal | Wrong value type for an existing column, malformed data, missing permission | -| `write-error`, `internal-error`, `dictionary-gap`, `cancelled`, `limit-exceeded`, `unknown` | Retriable | Disk pressure, a suspended table, a transient server fault | -| `not-writable` | Retriable on another endpoint | The server is a replica or cannot accept writes | -| `data-loss` | Abandoned | A corrupt store-and-forward journal was set aside | - -A retriable rejection is resent. For rejections that count toward the -poison-frame detector, if the same batch keeps being rejected after -`max_frame_rejections` (4) attempts spanning at least -`poison_min_escalation_window_millis` (5 minutes), the sender stops as for a -terminal error. The `dictionary-gap`, `unknown`, and `not-writable` categories -are exempt: they reset the poison episode instead of adding a strike. -Retriable rejections of symbol-dictionary catch-up frames are also exempt. -The six `on_*_error` connect-string keys are accepted but not applied by this -client. - -Handling notes: - -- **Message stability.** `serverMessage` is free-form English text from the - server. Its wording can change between releases: branch on `category`, not on - the text. -- **Sensitive data.** Server messages can contain column names and values. - Treat them as untrusted input, and redact them before sending them to - third-party error trackers or showing them to end users. -- **Correlation.** There is no server-side request ID. Correlate with the frame - sequence range, `tableName`, and `detectedAtMs`. - -#### Recovering from a terminal rejection - -After a terminal server rejection, the sender is permanently failed. An -already-pending `waitForAcknowledged()` for the rejected batch can reject with -`QwpIngressNackError`. Once the terminal failure is latched, new calls to -`waitForAcknowledged()`, `flush()`, or `close()` reject with -`QwpReplayRejectedError`, whose `status` and message repeat the server's. -Error handlers must allow either class depending on timing. Writing LONG arrays -with `longArrayColumn()` triggers a terminal rejection on every current server. - -Close the sender and create a new one. A pooled sender is replaced -automatically after the `close()` that reports the error. What happens to the -rejected batch depends on the mode: - -- **Without store-and-forward**, the failed sender's unacknowledged batches, - including the rejected one, are discarded with it, and so are rows that a - later borrower staged on it before the error surfaced. The new sender starts - empty. -- **With store-and-forward**, the rejected batch stays at the head of the - journal. Every new sender on that directory, including the pool's - replacement sender and the same client after a restart, sends it again and - fails the same way, with `QwpReplayRejectedError`. Pooled borrows keep - getting that journal, so ingestion through the client stops for every table, - not only the table in the rejected batch, until you act. Treat it as an - outage and alert on it from `onSenderError`. Fix the cause so that QuestDB - accepts the batch, for example by adjusting the table schema, or stop the - process and move the journal directory aside. Moving it aside discards every - unacknowledged batch in it, not only the rejected one. - -### Query errors - -Query errors reject both the `for await` iteration and `completion`: - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, - QwpEgressQueryTimeoutError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM no_such_table"); - for await (const batch of query) console.log(batch.rowCount); - await query.completion; - } catch (error) { - if (error instanceof QwpEgressQueryError) { - // Prints: 5 [14] table does not exist [table=no_such_table] - console.error(error.status, error.message); - } else if (error instanceof QwpEgressQueryTimeoutError) { - console.error("timed out"); - } else { - throw error; - } - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -`QwpEgressQueryError` has `status` (the QWP status code), `message` (the server -text, where parse errors start with the character position in brackets), and -`requestId` (a client-assigned `bigint` that numbers the queries of a -connection). The lease remains usable after a `QwpEgressQueryError`. - -| Status | Name | Meaning | -|---|---|---| -| `0x05` | PARSE_ERROR | SQL syntax error, unknown table or column, or a bind value the statement cannot use, such as a boolean for `LIMIT` | -| `0x06` | INTERNAL_ERROR | Execution failure, including a bind value that cannot be converted, such as `'abc'` compared with a DOUBLE column, and a column type that QWP cannot return | -| `0x08` | SECURITY_ERROR | Missing permission | -| `0x0a` | CANCELLED | The query was cancelled with `cancel()` | -| `0x0b` | LIMIT_EXCEEDED | A server limit was reached: the server-side query timeout, memory, or a result row too large to send | - -The `QWP_STATUS` export names these codes, for example -`QWP_STATUS.PARSE_ERROR`, so code can compare against constants instead of -numbers. A status alone does not separate a client mistake from a server -fault: `0x06` covers both bind values that cannot be converted and execution -failures, and `0x0b` covers both the server-side query timeout and memory -limits. - -Other query errors: - -| Error | Meaning | -|---|---| -| `QwpEgressQueryTimeoutError` | The query deadline expired and cancellation started. Has `requestId` and `timeoutMs`. | -| `QwpEgressQueryAbandonedError` | Iteration ended early, for example with `break`. | -| `QwpEgressQueryCancelTimeoutError` | QuestDB did not confirm a cancellation in time; the connection was closed. | -| `QwpEgressSessionClosedError` | The query connection is closed. | -| `QwpReconnectExhaustedError` | Failover gave up; see [Query failover](#query-failover). | - -As with ingestion, the message text is not stable, may echo parts of the SQL, -and has no server-side correlation ID beyond `requestId`. - -### Connection-level errors - -| Error | Raised when | -|---|---| -| `QwpUpgradeError` | Connecting to an endpoint failed. `kind` is `authentication` (HTTP 401 or 403), `role-rejected`, `http-rejected`, `version-mismatch`, `capability-mismatch`, `timeout`, or `transport`. It also carries `statusCode`, `retryable`, and `url`. | -| `QwpFailoverError` | Every endpoint in a multi-host list failed. `attempts` holds each endpoint and its error. | -| `QwpPoolResourceError` | The pool could not open a new connection. `cause` holds the error above. | -| `QwpPoolAcquireTimeoutError` | Every pooled connection stayed leased past `acquire_timeout_ms`. | -| `QwpReconnectExhaustedError` | The reconnect budget ran out. The sender or query failed permanently. | -| `QwpRoleMismatchError` | No endpoint has the role that `target` requires. | -| `QwpDurableAckUnavailableError` | `request_durable_ack=on`, but the server does not support it. | -| `QwpClientClosedError` | The pooled client, or a returned lease, is already closed. | - -The pooled client wraps every failure to open a connection, from -`connectQwpNodeClient()`, `db.connect()`, `borrowSender()`, or -`borrowQuery()`, in a `QwpPoolResourceError`. Unwrap its `cause` before -checking for a specific error. When `addr` lists several hosts, the cause is a -`QwpFailoverError` whose `attempts` hold the error of each endpoint. When -initial-connect retry is on, for example with a `failover_*` or `reconnect_*` -key, or with a typed `egressSession.reconnect` object for query connections -(see [Typed reconnect policy](#typed-reconnect-policy)), the cause is a -`QwpReconnectExhaustedError` instead, and its own `cause` holds the last -attempt's error: - -```typescript -import { - connectQwpNodeClient, - QwpFailoverError, - QwpPoolResourceError, - QwpReconnectExhaustedError, - QwpUpgradeError, -} from "@questdb/nodejs-client"; - -// The errors behind a failed connection, one per endpoint tried. -function connectionErrors(error: unknown): unknown[] { - let cause = error instanceof QwpPoolResourceError ? error.cause : error; - // With initial-connect retry on, the last attempt's error is wrapped. - if (cause instanceof QwpReconnectExhaustedError) cause = cause.cause; - return cause instanceof QwpFailoverError - ? cause.attempts.map((attempt) => attempt.error) - : [cause]; -} - -try { - const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); - await db.close(); -} catch (error) { - for (const cause of connectionErrors(error)) { - if (cause instanceof QwpUpgradeError && cause.kind === "authentication") { - console.error("QuestDB rejected the credentials:", cause.message); - } else { - console.error("cannot connect:", cause); - } - } - throw error; -} -``` - -An authentication rejection (HTTP 401 or 403) is terminal before a sender's -first successful connection and for query connections. It stops the endpoint -walk because credentials are assumed to be shared across the cluster. - -After a successful connection, regular senders with `sf_dir` or in background -memory mode (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying -authentication rejections indefinitely. This lets buffered data drain once -server-side authentication is restored. Senders in default memory mode, and -orphan drainers, do not have this exception. See -[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) -for how other clients behave. - -Endpoints in error messages have any embedded credentials removed. - -#### Connection timeouts - -Two transport deadlines bound WebSocket setup: `connect_timeout` covers DNS -and the TCP/TLS connection, and `auth_timeout_ms` covers the upgrade and -authentication. Both default to 15 seconds, and `auth_timeout_ms` inherits -`connect_timeout` when only the latter is set. A timeout in either phase -produces a `QwpUpgradeError` whose `timeoutPhase` is `connect` or -`authentication`. - -After the upgrade, a query connection has a separate 5-second deadline for -the initial QWP `SERVER_INFO` frame. Configure it with the typed option -`egressSession.serverInfoTimeoutMs`; raising the transport deadlines does not -change it. Expiry produces an ordinary `Error` with the message -`timed out waiting for QWP SERVER_INFO`, not a `QwpUpgradeError`. - -### Logging - -The client writes its own messages to the console by default, at the `error`, -`warn`, and `info` levels. To route a sender's messages, such as warnings about -rows discarded on close, through your logger, pass a `QwpSenderLogger` -function: `{ sender: { log } }` as the second argument of -`connectQwpNodeClient()`, or `{ log }` for `Sender.fromConfig()`. Its signature -is `(level: "error" | "warn" | "info" | "debug", message: string | Error)`, so -convert `message` with `String()` if your logger takes strings only. The -function also receives `debug` messages, one per staged row, so filter by -level. - -Rejected batches and session errors go to `ingressSession.onSenderError` and -`ingressSession.onError`. Their defaults log to the console, so replace both to -route them through your logger. Some messages from other parts of the client, -such as store-and-forward recovery, always go to the console. - -## Failover and high availability - -:::note Enterprise - -Failing over between several QuestDB hosts requires QuestDB Enterprise -replication. Reconnecting to a single restarted server works in open source -too. - -::: - -### Multiple endpoints - -List several hosts in `addr`: - -```text -wss::addr=db-a.example.com:9000,db-b.example.com:9000,db-c.example.com:9000; -``` - -The client ranks endpoints by observed health and by `zone`, and on a -connection loss moves to the next usable one. `addr` is shared by ingestion and -queries. - -Ingestion always needs the primary: replicas refuse writes, and the sender -walks the list until it finds the current primary. Queries can use any node. -`target` selects which roles queries accept: `any` (the default), `primary`, or -`replica`. Set it with the typed `egress` option, as below, because in the -connect string `target` also filters ingestion (see the caution that follows). - -`target` is a strict filter, not a preference. With `replica`, queries never -fall back to the primary, and they fail when no replica is reachable, including -against a single open source server. Because the pooled client opens a query -connection at startup, `connectQwpNodeClient()` then fails too, with a -`QwpPoolResourceError` whose `cause` leads to a `QwpRoleMismatchError`; see -[Connection-level errors](#connection-level-errors) to unwrap it. To start -without a replica, also set `query_pool_min=0`. Queries borrowed before a -replica is reachable then reject with `QwpPoolResourceError`: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -// Queries run on replicas only. Ingestion still follows the primary. -const db = await connectQwpNodeClient( - "wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN;" + - // Start, and ingest, even while no replica is reachable. - "query_pool_min=0;", - { egress: { target: "replica" } }, -); -await db.close(); -``` - -`zone` prefers endpoints in the same zone. - -:::caution `target` in the connect string also filters ingestion - -Unlike the Java client, the Node.js client applies `target` and `zone` from -the connect string to ingestion as well as queries. `target=replica` in the -connect string therefore stops ingestion from reaching the primary. To read -from replicas and write to the primary with one client, keep `target` out of -the connect string and set it for queries only: -`connectQwpNodeClient(conf, { egress: { target: "replica" } })`. - -::: - -### Ingestion reconnect - -When the connection drops, the sender reconnects with exponential backoff and -jitter, then resends every unacknowledged batch: - -| Key | Default | Purpose | -|---|---|---| -| `reconnect_initial_backoff_millis` | `100` | First retry delay. | -| `reconnect_max_backoff_millis` | `5000` | Longest delay between retries. | -| `reconnect_max_duration_millis` | `300000` (5 minutes) | Budget for one outage in default memory mode. `0` removes the limit. | -| `initial_connect_retry` | `off` | Whether the first connection retries: `off` fails fast, `on` (or `sync`) retries within the budget, `async` connects in the background. | - -Whether the sender gives up depends on the [ingestion mode](/docs/connect/clients/nodejs/#ingestion-modes): - -- **Default memory mode** retries for up to `reconnect_max_duration_millis` - per outage. When the budget runs out, the sender fails permanently with - `QwpReconnectExhaustedError`, and its unsent rows are lost. The Java - reference client retries indefinitely in this mode instead. -- **Background memory mode** (`initial_connect_retry=async` or - `lazy_connect=on`) and **store-and-forward** (`sf_dir`) retry indefinitely. - -Setting any `reconnect_*` key also makes a sender's first connection retry -within the budget, as if `initial_connect_retry=on`. Set -`initial_connect_retry=off` explicitly to keep a fail-fast start. The keys do -not apply to query connections: the pooled client still opens its query pool -at startup, so `connectQwpNodeClient()` fails fast while QuestDB is down unless -you also set `query_pool_min=0` or enable query retries (see -[Connection events](#connection-events)). - -Replay after a reconnect is at least once: a batch that QuestDB committed just -before the connection dropped is sent again. Write to a deduplicated table, as -described under [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward), to keep replayed rows -from inserting duplicates. - -### Query failover - -If the connection fails during a query, the client reconnects, to another -endpoint when there is one, and runs the query again from the start: - -| Key | Default | Purpose | -|---|---|---| -| `failover` | `on` | Set `off` to fail the query instead of retrying. | -| `failover_max_attempts` | `8` | Connection attempts per failure. Each attempt tries every endpoint in `addr`. | -| `failover_backoff_initial_ms` | `50` | First retry delay. | -| `failover_backoff_max_ms` | `1000` | Longest delay between retries. | -| `failover_max_duration_ms` | `30000` | Time budget per failure. | - -The attempt limit and the time budget apply together, and whichever is reached -first ends the failover. When attempts fail fast, for example with connection -refused while a server restarts, the 8 attempts and their backoff of 50 ms to -1 second, with jitter, take only about 1 to 3 seconds, and at most about 4.5 -seconds, long before the 30-second budget. To ride out a longer restart, raise `failover_max_attempts`, -or set `maxAttempts: 0` in a typed `egressSession.reconnect` object to remove -the attempt limit and rely on the time budget alone; see -[Typed reconnect policy](#typed-reconnect-policy). - -When failover gives up, the query rejects with `QwpReconnectExhaustedError`, -and the lease stays failed even after QuestDB recovers: close it and borrow a -new one. A `QwpEgressQueryError` from the server is a query result and never -triggers failover. Replaying an in-flight `query()` also re-executes DDL and -DML: an `INSERT` may run twice if its completion was lost. For non-idempotent -SQL, use a separate client configured with `failover=off` and check an -uncertain outcome before retrying; see -[DDL and DML statements](/docs/connect/clients/nodejs/#ddl-and-dml-statements). - -:::warning Clear partial results when a query restarts - -A re-executed query starts again from the first row. Batches that were queued -but not yet consumed are discarded for you, but rows your loop already -processed are delivered again. If your code accumulates rows, clear them when -the query restarts; otherwise it sees the first part of the result twice. - -::: - -Every batch has a `batchSequence` starting at `0n`, including the first batch -after a replay. Clear accumulated rows on that batch. A replay may return -**zero rows and no batches**, though, leaving prior rows in your accumulator. -`egressSession.onReplayReset` also clears it when that happens. Since this -callback is shared across the pool and request IDs are per connection, the -example limits the pool to one active query: - -```typescript -import { - connectQwpNodeClient, - QwpReconnectExhaustedError, -} from "@questdb/nodejs-client"; - -const rows: (readonly unknown[])[] = []; -const db = await connectQwpNodeClient( - "ws::addr=db-a.example.com:9000,db-b.example.com:9000;query_pool_max=1;", - { - egressSession: { - onReplayReset: () => { - rows.length = 0; - }, - }, - }, -); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM trades LIMIT 100000", { - // The deadline covers the whole query, including a re-execution. - timeoutMs: 30_000, - }); - for await (const batch of query) { - // Sequence 0 starts the result, both initially and after a failover. - if (batch.batchSequence === 0n) rows.length = 0; - for (const row of batch.rows()) rows.push(row); - } - await query.completion; - console.log(`${rows.length} rows`); - } catch (error) { - if (!(error instanceof QwpReconnectExhaustedError)) throw error; - // Failover gave up. This lease stays failed: return it, retry later. - console.error("no endpoint could run the query:", error.message); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -The reset needs the rows of the current attempt in one place you can discard. -When your code passes rows on as they arrive, for example by streaming them to -an HTTP response, it cannot take them back after a restart. Then choose one -of these instead: - -- Run such queries on a separate client with `failover=off`, so that a lost - connection fails the query instead of restarting it, and retry the whole - request. -- Keep each result small enough to buffer, for example by paging with - `WHERE timestamp < $1 ORDER BY timestamp DESC LIMIT n`, binding the oldest - timestamp of the previous page. -- Treat `batchSequence === 0n` after rows have left as an error, and abort the - downstream response instead of sending duplicates. - -`egressSession.onReplayReset` runs before a query is replayed, including when -that replay returns no batches. Its event has `requestId`, `endpoint`, -`previousEndpoint`, `serverInfo`, and `cause`. The `requestId` matches -`query.requestId`, but request IDs are numbered per connection and every lease -of a pooled client shares the callback: it cannot identify which of several -concurrent queries restarted. Use a dedicated, single-query client when the -callback resets result state, as above. For concurrent results that cannot be -isolated, set `failover=off` and retry the whole query after a transport error; -`batchSequence === 0n` alone cannot detect a zero-batch replay. - -### Typed reconnect policy - -Reconnect and failover behavior comes from the connect-string keys above, or -from typed `reconnect` objects in the second argument of -`connectQwpNodeClient()`: `ingressSession.reconnect` for senders and -`egressSession.reconnect` for queries. You need the object to register -`onEvent` for [connection events](#connection-events). - -:::caution A typed `reconnect` object replaces the connect-string keys - -`ingressSession.reconnect` (or `qwp.session.reconnect` on a `Sender`) replaces -the whole reconnect policy parsed from the `reconnect_*`, -`max_frame_rejections`, and `poison_min_escalation_window_millis` keys, and -`egressSession.reconnect` replaces the policy parsed from `failover*` keys. -Fields you leave out of the object take the built-in defaults, not the values -from the connect string. When you supply the object, for example to register -`onEvent`, set every bound you rely on in it. - -::: - -The object's fields and the connect-string keys they replace: - -| Field | Ingestion key, default | Query key, default | -|---|---|---| -| `maxAttempts` | None, `0` (unlimited) | `failover_max_attempts`, `8`. The key accepts `1` or more; the typed field also accepts `0`, unlimited | -| `initialBackoffMs` | `reconnect_initial_backoff_millis`, `100` | `failover_backoff_initial_ms`, `50` | -| `maxBackoffMs` | `reconnect_max_backoff_millis`, `5000` | `failover_backoff_max_ms`, `1000` | -| `maxDurationMs` | `reconnect_max_duration_millis`, `300000` | `failover_max_duration_ms`, `30000` | -| `maxFrameRejections` | `max_frame_rejections`, `4` | Not used | -| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis`, `300000` | Not used | -| `onEvent` | None | None | - -`egressSession: { reconnect: false }` is the typed equivalent of -`failover=off`. Senders in background memory mode or with `sf_dir` retry -indefinitely: `maxAttempts` and `maxDurationMs` do not end their retries, so -the full example's `reconnect: { onEvent }` with `sf_dir` keeps retrying -through an outage of any length. - -The two directions treat the first connection differently: - -- **Senders**: setting any `reconnect_*` key makes the first connection retry - within the budget, as if `initial_connect_retry=on`. A typed - `ingressSession.reconnect` object does not, so the first connection still - fails fast. Set `initial_connect_retry` in the connect string to choose the - startup behavior. -- **Query connections**: the first connection retries within the failover - budget, for retryable errors, when you supply an `egressSession.reconnect` - object, set `failover=on` explicitly, or set a `failover_*` key without - `failover=off`. Otherwise it is attempted once. `failover=off` and - `egressSession.reconnect: false` turn reconnects off entirely. - -### Connection events - -Register `reconnect.onEvent` to observe connections. Events are delivered -asynchronously through a bounded queue (64 by default). The connect-string -`connection_listener_inbox_capacity` key configures the ingestion queue only; -for query events, set the typed `egressSession.connectionListenerInboxCapacity` -option. When a queue overflows, its oldest events are dropped and counted in -the metrics. - -```typescript -import { - connectQwpNodeClient, - QWP_RECONNECT_EVENT_KIND, - type QwpReconnectEvent, -} from "@questdb/nodejs-client"; - -function onEvent(event: QwpReconnectEvent) { - switch (event.kind) { - case QWP_RECONNECT_EVENT_KIND.RECONNECTING: - console.warn("connection lost, reconnecting:", event.cause); - break; - case QWP_RECONNECT_EVENT_KIND.FAILED_OVER: - console.warn(`failed over to ${String(event.endpoint)}`); - break; - default: - console.info(event.kind, String(event.endpoint ?? "")); - } -} - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - // Each object replaces the reconnect_* or failover* keys from the connect - // string. Fields left out use the built-in defaults. - ingressSession: { - reconnect: { onEvent }, - // No event marks a terminal failure: it arrives here instead. - onError: (event) => { - if (event.terminal) console.error("ingestion stopped:", event.error); - }, - }, - egressSession: { reconnect: { onEvent } }, -}); -await db.close(); -``` - -Supplying the `reconnect` objects also changes how the first connection is -retried; see [Typed reconnect policy](#typed-reconnect-policy). - -| Kind | Meaning | -|---|---| -| `connected` | The first connection succeeded. | -| `reconnecting` | The active connection was lost. `cause` holds the error. | -| `attempt-failed` | One connection attempt failed. The client retries if the error is retryable and its budget allows; otherwise this is the last event before the failure is reported. | -| `reconnected` | Reconnected to the same endpoint. | -| `failed-over` | Reconnected to a different endpoint. `previousEndpoint` holds the old one. | -| `durable-ack-unavailable` | A sender is waiting for an endpoint that supports durable acknowledgement. A background-started sender emits this from startup, including with `sf_dir`; a foreground-started sender with `sf_dir` retries after its first successful connection. | -| `durable-ack-persistent-failure` | An orphan drainer gave up waiting for durable acknowledgement support. | -| `primary-unavailable` | An orphan drainer, which recovers a journal left by another sender (see [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward)), found no endpoint that currently accepts writes. It keeps retrying. Regular senders do not emit it. | - -The `QWP_RECONNECT_EVENT_KIND` constants name these kinds: `CONNECTED`, -`RECONNECTING`, `ATTEMPT_FAILED`, `RECONNECTED`, `FAILED_OVER`, -`DURABLE_ACK_UNAVAILABLE`, `DURABLE_ACK_PERSISTENT_FAILURE`, and -`PRIMARY_UNAVAILABLE`. - -`reconnected` and `failed-over` are mutually exclusive: code that tracks the -current node must handle both. The client has no property that says whether it -is connected right now. To report it, for example in a health check, track the -latest event: after `reconnecting` the connection is down, and `connected`, -`reconnected`, or `failed-over` mean it is up. A query that reconnects runs -again from its first batch; see -[Query failover](#query-failover) for resetting accumulated rows. - -No event marks a terminal failure. When a sender stops retrying, because its -reconnect budget ran out or the error cannot be retried, -`ingressSession.onError` receives a `QwpIngressErrorEvent` with -`terminal: true`, even while the sender is idle. The event also has `error`, -`timestampMs`, and, for a server rejection, `senderError`. The sender's next `flush()`, auto-flushing `at()`, or -`close()` then rejects with the same error, such as -`QwpReconnectExhaustedError`. A query that cannot fail over rejects its -iteration and `completion` instead. - -For ingestion, `ingressSession` also accepts `onProgress`, for published, -acknowledged, and durably acknowledged sequences, and `onError`, for session -errors. `sender.metrics` returns a snapshot of the sender's counters, including -`metrics.ingress` with the replay queue, reconnect, and notification counters. - - - -## Configuration reference - -The [connect string reference](/docs/connect/clients/connect-string/) documents -every key. The Node.js client's defaults and deviations: - -| Key | Default | Notes | -|---|---|---| -| `addr` | required | Comma-separated or repeated for failover. Port defaults to `9000`. | -| `username`, `password`, `token` | none | Basic or bearer authentication. | -| `tls_verify`, `tls_roots` | `on`, Node.js CA bundle | `wss` only. `tls_roots` must be PEM. `tls_roots_password` is rejected. | -| `connect_timeout`, `auth_timeout_ms` | `15000` | DNS and TCP/TLS connection, and upgrade deadlines, in milliseconds. See [Connection timeouts](#connection-timeouts). | -| `auto_flush` | `on` | Master switch for the three triggers. | -| `auto_flush_rows` | `1000` | `0` disables. `off` is rejected. | -| `auto_flush_interval` | `100` | Milliseconds. `0` disables. `off` is rejected. | -| `auto_flush_bytes` | disabled | Size, or `off`. | -| `close_flush_timeout_millis` | `5000` | ACK wait in a standalone sender's `close()`. | -| `transaction` | `off` | Keep auto-flushed batches in an open transaction until `flush()`. | -| `request_durable_ack`, `durable_ack_keepalive_interval_millis` | `off`, `200` | Enterprise. Explicitly setting the keepalive interval alone requests durable ACK; a negative interval is rejected. | -| `max_name_len` | `127` | Maximum table and column name length, in UTF-8 bytes. | -| `reconnect_initial_backoff_millis`, `reconnect_max_backoff_millis` | `100`, `5000` | Ingestion reconnect backoff. | -| `reconnect_max_duration_millis` | `300000` | Ingestion budget per outage in default memory mode. `0` removes it. | -| `max_frame_rejections`, `poison_min_escalation_window_millis` | `4`, `300000` | Poison-frame detector: rejections of one batch, and the minimum time they must span, before the sender stops. | -| `initial_connect_retry` | `off` | `off`, `on`/`sync`, or `async`. | -| `sf_dir`, `sender_id` | none, `default` | Store-and-forward journal location. | -| `sf_durability` | `memory` | `memory`, `periodic`, or `append`. Requires `sf_dir` when explicitly set, even to `memory`. | -| `sf_max_total_bytes` | `10g` with `sf_dir`, `128m` without | [Journal size target](/docs/connect/clients/nodejs/#sf-capacity), not a hard disk limit; memory queue cap without `sf_dir`. | -| `sf_max_segment_bytes` | `4m` with `sf_dir`, none without | Journal segment size, which also caps a batch. | -| `sf_append_deadline_millis` | `30000` | How long a full journal or queue blocks publishing. | -| `drain_orphans`, `max_background_drainers` | `off`, `4` | Adopt journals left by crashed processes. Both require `sf_dir` when explicitly set. | -| `target`, `zone` | `any`, none | Endpoint role and zone preference. Apply to ingestion too. | -| `failover`, `failover_max_attempts`, `failover_max_duration_ms` | `on`, `8`, `30000` | Query failover. | -| `compression`, `compression_level` | `raw`, `1` | Query result compression. Explicit `compression_level` requires `compression=zstd` or `auto`. | -| `initial_credit`, `buffer_pool_size`, `max_batch_rows` | `0`, `4`, server default | Query flow control. | -| `client_id` | `typescript/` | Sent to the server for diagnostics. | -| `error_inbox_capacity`, `connection_listener_inbox_capacity` | `256`, `64` | Queues for rejection callbacks and ingestion connection events. Query event queue: typed `egressSession.connectionListenerInboxCapacity`. | -| Pool keys | see [Pool settings](#pool-settings) | Applied by the pooled client. A standalone `Sender` also applies `lazy_connect`. | - -The -[API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) -covers every type and option. The -[QWP guide](https://github.com/questdb/nodejs-questdb-client/blob/main/QWP.md) -in the client repository describes the delivery semantics in depth. - -### Programmatic options - -Callbacks, custom agents, and other settings a string cannot express go in the -second argument, a `QwpNodeClientConfigOptions` object. When the connect string -and typed options set the same option, the typed value wins. Credentials and -TLS are the exception: a typed `webSocket.authorization` header cannot be -combined with `token`, `username`, or `password` in the string, and a typed -`webSocket.agent` cannot be combined with `tls_verify` or `tls_roots`. Both -combinations are rejected. - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - // QwpSenderOptions for every pooled sender - sender: { awaitServerAck: true }, - // Ingestion callbacks and replay settings - ingressSession: { - onSenderError: (error) => - console.error("rejected batch", error.category, error.serverMessage), - }, - // Query session defaults - egressSession: { - queryTimeoutMs: 30_000, - cancelDrainTimeoutMs: 5_000, - serverInfoTimeoutMs: 10_000, - }, - // Egress-only routing and compression - egress: { compression: "zstd" }, - // Pool sizes and timeouts - pool: { senderPoolMax: 2, queryPoolMax: 8 }, -}); -await db.close(); -``` - -The typed `egress` section takes `target`, `zone`, `compression`, -`compressionLevel`, and `maxBatchRows`. The other sections are `webSocket` -(connection settings shared by both directions, such as `agent` or -`connectTimeoutMs`) and `storeAndForward`, the journal settings described -under [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward). Its fields match the connect -string keys: - -| Typed field | Connect string key | -|---|---| -| `directory` | `sf_dir` | -| `maxBytes` | `sf_max_total_bytes` | -| `maxSegmentBytes` | `sf_max_segment_bytes` | -| `durability` | `sf_durability` | -| `checkpointIntervalMs` | `sf_sync_interval_millis` | -| `appendDeadlineMs` | `sf_append_deadline_millis` | -| `drainOrphans` | `drain_orphans` | -| `maxBackgroundDrainers` | `max_background_drainers` | - -The second argument has no field for `sender_id`; set it in the connect -string. - -`Sender.fromConfig()` takes `{ log, agent, qwp }` as its second argument, -where `qwp` has the sections `webSocket`, `session` (the equivalent of -`ingressSession`), `sender`, and `udp`. - -The `reconnect` objects in `ingressSession` and `egressSession` replace the -whole reconnect policy parsed from the connect string; see -[Typed reconnect policy](#typed-reconnect-policy) before you set one. - -### Differences from other clients - -The Node.js client differs from the Java reference client, and from the shared -[connect string reference](/docs/connect/clients/connect-string/), in these -places: - -| Area | Node.js behavior | -|---|---| -| Outage budget | A sender in default memory mode gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`. Senders in background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) or with `sf_dir` retry indefinitely. See [Ingestion reconnect](#ingestion-reconnect). | -| `target` and `zone` | Also apply to ingestion. Set a query-only role with the typed `egress.target` option. See [Multiple endpoints](#multiple-endpoints). | -| Authentication rejected after a first connection | Senders with `sf_dir` or in background memory mode keep retrying. Other senders and query connections fail. See [Connection-level errors](#connection-level-errors). | -| Durable acknowledgement unavailable | Senders started in the background retry from startup even with `sf_dir`; foreground store-and-forward senders fail on first connect, but retry after a successful connection. They emit `durable-ack-unavailable` while retrying. See [Durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). | -| `sf_durability` and SF-only keys | Also accepts `append`, but rejects explicit `sf_durability` (even `memory`), `sf_sync_interval_millis`, `drain_orphans` (even `off`), `max_background_drainers`, or `catch_up_cap_gap_min_escalation_window_millis` without `sf_dir`. Unlike Java, do not pass these keys in a memory-mode shared string. | -| `sf_max_total_bytes` with `sf_dir` | A journal size target that can be exceeded, not a hard limit. See [Journal capacity](/docs/connect/clients/nodejs/#sf-capacity). | -| `sf_dir` path creation | Creates missing parent directories and the slot recursively; Java and Rust-derived clients only create `sf_dir` and the slot. | -| `durable_ack_keepalive_interval_millis` | Explicitly setting this key, even to `0`, also requests durable ACK; it fails against OSS if the sender connects. Negative values throw `RangeError` (the shared reference treats them as disabled). | -| Journal lock | A `.lock.owner` directory that can outlive a crashed process and that other clients' operating-system locks do not see. See [Lock recovery](/docs/connect/clients/nodejs/#sf-lock-recovery). | -| `max_lifetime_ms` | Closes idle connections above the pool minimum only. Connections at the minimum are not recycled. | -| Connect string parsing | `0`, not `off`, disables `auto_flush_rows` and `auto_flush_interval`, and the interval runs from the last flush or from sender creation. Size values take single-letter suffixes only. `compression_level` requires `compression=zstd` or `auto`. `tls_roots` must be PEM; `tls_roots_password`, `init_buf_size`, and `max_buf_size` are rejected. | -| Initial connection and reconnect | `lazy_connect=on` opens senders in the background at startup instead of waiting for a borrow. A query's first connection retries with explicit `failover=on`, a `failover_*` key (unless `failover=off`), or typed `egressSession.reconnect`, not just from the default `failover=on`. See [Starting while QuestDB is down](#starting-while-questdb-is-down) and [Typed reconnect policy](#typed-reconnect-policy). Ingestion reconnect uses full jitter (delay from 0 up to the backoff ceiling), not the equal-jitter schedule in the shared failover guide. | -| `connect_timeout` | Also covers DNS and the TLS handshake, and `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. See [Connection timeouts](#connection-timeouts). | -| `tls_roots` default | The CA certificates bundled with Node.js, not the operating system's trust store. See [TLS](/docs/connect/clients/nodejs/#tls). | -| Defaults | `connect_timeout` is `15000` and `poison_min_escalation_window_millis` is `300000`. `close_flush_timeout_millis` is `5000`, as in the Rust, C, C++, Python, and Go clients; Java and .NET use `60000`. | -| Close after an ACK timeout | A standalone sender's `close()` rejects with `QwpSenderCloseTimeoutError` instead of logging a warning. The pooled client's `db.close()` resolves and reports the timeout, best-effort, to `ingressSession.onError`. See [Closing a sender](/docs/connect/clients/nodejs/#closing-a-sender). | -| Error reports | Categories and policies are lowercase, hyphenated strings, such as `schema-mismatch` and `retriable-other`. See [Ingestion errors](#ingestion-errors). | -| Pool and query keys on a standalone `Sender` | The `Sender` logs a warning for the pool and query-only keys it ignores. It applies `client_id` and `lazy_connect`. | -| `connection_listener_inbox_capacity` | Sets the ingestion event inbox only. Set the query inbox with typed `egressSession.connectionListenerInboxCapacity`; see [Connection events](#connection-events). | -| `on_*_error` keys | Accepted but not applied. | - -## Migration - -### From ILP to QWP - -The row API is unchanged, so existing `Sender` code migrates by changing the -connect string and calling `connect()`: - -```diff -- const sender = await Sender.fromConfig("http::addr=localhost:9000"); -+ const sender = await Sender.fromConfig("ws::addr=localhost:9000"); -+ await sender.connect(); -``` - -| Aspect | ILP over HTTP | QWP over WebSocket | -|---|---|---| -| Connect string schema | `http::`, `https::` | `ws::`, `wss::` | -| Auto-flush rows | 75,000 (600 over TCP) | 1,000 | -| Auto-flush interval | 1,000 ms | 100 ms | -| `flush()` completes when | QuestDB responds to the HTTP request | The batch is published; the ACK arrives later | -| Server rejection | `flush()` throws | Asynchronous: `onSenderError`, `waitForAcknowledged()`, or `flush()` with `awaitServerAck` | -| Rows staged at `close()` | Lost unless flushed | Published; waits up to 5 seconds for ACK, then unacknowledged rows may be lost without `sf_dir` | -| Reconnect and replay | Retries one request for `retry_timeout` | Automatic, with replay of unacknowledged batches | -| Store-and-forward, querying, pooling | Not available | Available | -| Column types | ILP types | More types, subject to [column-method](/docs/connect/clients/nodejs/#column-methods) and [array](/docs/connect/clients/nodejs/#arrays) support | - -Legacy keys such as `retry_timeout`, `request_timeout`, `init_buf_size`, -`max_buf_size`, `protocol_version`, and `tls_ca` are rejected on `ws`/`wss` -with a hint. It names the replacement where there is one -(`retry_timeout` becomes `reconnect_max_duration_millis`, and `tls_ca` becomes -`tls_roots`), and otherwise says that the key applies only to ILP or that QWP -negotiates the setting itself. To keep ILP-sized batches, set -`auto_flush_rows` and `auto_flush_interval` explicitly. Migrate one sender at a -time: ILP and QWP senders can run side by side. - -### Upgrading from 4.x - -Version 5.0.0 keeps the ILP API and adds QWP. Changes that affect existing ILP -code: - -- **Null values.** Passing `null` or `undefined` to a column or symbol method now - omits the column. Existing nullable columns store NULL; BOOLEAN defaults to - `false`, and BYTE and SHORT default to `0` (see [Null values](/docs/connect/clients/nodejs/#null-values)). - Earlier versions threw a type error for most such values. Validate data - before calling the sender if you relied on the error. -- **Decimal scale.** `decimalColumn()` over ILP rejects a non-integer `scale` - with a `RangeError`. Earlier versions silently coerced it, writing `2.5` as - scale 2 and `NaN` as scale 0. -- **`intColumn()`** also accepts a `bigint`, for LONG values beyond - `Number.MAX_SAFE_INTEGER`. -- **TCP authentication** now works on Node.js 26, which rejects the JWK the - client previously built. -- **New dependency.** The package now depends on `ws`, used for QWP. - -## ILP transports (legacy) - -The Node.js `Sender` still ingests over ILP, for existing deployments and for -servers without QWP. To move ILP code to QWP, see -[From ILP to QWP](#from-ilp-to-qwp); for behavior changes in 5.0.0, see -[Upgrading from 4.x](#upgrading-from-4x). ILP senders support HTTP (`http::`, -`https::`) and TCP (`tcp::`, `tcps::`) transports: - -```typescript -import { Sender } from "@questdb/nodejs-client"; - -const sender = await Sender.fromConfig( - "http::addr=localhost:9000;username=admin;password=quest;", -); -try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.00044) - .at(Date.now(), "ms"); - // ILP does not flush on close: rows still buffered at close() are lost. - await sender.flush(); -} finally { - await sender.close(); -} -``` - -- HTTP connects per request, so `connect()` is not needed; TCP transports - require `await sender.connect()`. `token=...` selects bearer authentication - over HTTP. Over TCP, `username` and `token` set the JWK key ID and private - key. -- Over HTTP, `flush()` sends the buffer as one request and throws if QuestDB - rejects it. Data is transactional only for a single-table request. A - multi-table request can commit earlier tables before a later table fails, - so a failed flush does not mean no data was committed. Schema changes, such - as automatically added columns, are not rolled back even for a single-table - request. See [HTTP transaction semantics](/docs/connect/compatibility/ilp/overview/#http-transaction-semantics). -- Decimals need ILP protocol version 3: HTTP negotiates it automatically, and - TCP needs `protocol_version=3`. Arrays need version 2 or later. -- Undici is the default HTTP agent. Set `stdlib_http=on` to use the Node.js - `http` module instead. - -For ILP options, see the -[`SenderOptions` reference](https://questdb.github.io/nodejs-questdb-client/classes/_questdb_nodejs-client.SenderOptions.html) -and the [ILP overview](/docs/connect/compatibility/ilp/overview/). - -## Full example: Ingestion and querying with failover - -A production-oriented pattern that ingests trades and queries recent prices, -with TLS, a token, several hosts, error handling, and failover handling. Before -running it, create the deduplicated table on the primary (or reuse the table -from [Store-and-forward](/docs/connect/clients/nodejs/#store-and-forward)). If the table is missing, QWP -creates it without deduplication: - -```questdb-sql -CREATE TABLE IF NOT EXISTS trades_sf ( - timestamp TIMESTAMP, - trade_id VARCHAR, - symbol SYMBOL, - side SYMBOL, - price DOUBLE, - amount DOUBLE -) TIMESTAMP(timestamp) PARTITION BY DAY -DEDUP UPSERT KEYS(timestamp, trade_id); -``` - -Replace the sample events with source-assigned trade IDs and timestamps. Keep -both values unchanged when retrying the same event, and use a writable, -persistent `sf_dir` so unacknowledged rows survive a shutdown: - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, - QwpIngressAckTimeoutError, - QwpPoolResourceError, - QWP_RECONNECT_EVENT_KIND, - QWP_SENDER_ERROR_POLICY, - type QwpReconnectEvent, - type QwpSenderError, -} from "@questdb/nodejs-client"; - -const token = process.env.QDB_TOKEN; -if (!token) throw new Error("QDB_TOKEN is not set"); - -// This example runs one query at a time so its replay callback can clear it. -const recentPrices: (readonly unknown[])[] = []; - -// Replace with your alerting. -function alertOperator(message: string) { - console.error("ALERT:", message); -} - -function logConnection(event: QwpReconnectEvent) { - if (event.kind !== QWP_RECONNECT_EVENT_KIND.ATTEMPT_FAILED) { - const endpoint = String(event.endpoint ?? ""); - console.info("questdb connection:", event.kind, endpoint); - } -} - -const db = await connectQwpNodeClient( - "wss::addr=db-primary.example.com:9000,db-replica.example.com:9000;" + - `token=${token};` + - // append: every flush waits for a disk sync; see "Store-and-forward". - "sf_dir=/var/lib/my-service/qdb-sf;sender_id=trade-service;" + - // Limit offline batches below the default server's 2 MiB limit. - "sf_durability=append;sf_max_segment_bytes=1m;sender_pool_max=4;" + - // Query pool stays cold while replicas are down; one active query so the - // replay callback below can reset its state even if replay has no batches. - "query_pool_min=0;query_pool_max=1;", - { - // Queries run on replicas only, never on the primary; ingestion always - // follows the primary. - egress: { target: "replica", compression: "zstd" }, - ingressSession: { - onSenderError: (error: QwpSenderError) => { - console.error("batch rejected:", error.category, error.serverMessage); - // A terminally rejected batch stays in the journal and blocks - // ingestion through this client, for every table, until it is fixed. - if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { - alertOperator(`QuestDB rejected a batch: ${error.serverMessage}`); - } - }, - // Terminal failures, such as a batch that QuestDB rejects terminally. - onError: (event) => { - if (event.terminal) console.error("ingestion stopped:", event.error); - }, - // Replaces any reconnect_* keys; omitted fields use the defaults. - reconnect: { onEvent: logConnection }, - }, - egressSession: { - queryTimeoutMs: 30_000, - // Replaces any failover* keys; omitted fields use the defaults. - // No 8-attempt limit (maxAttempts 0): failover can last up to 30 s. - reconnect: { - maxAttempts: 0, - maxDurationMs: 30_000, - onEvent: logConnection, - }, - onReplayReset: (event) => { - recentPrices.length = 0; - console.warn("query restarts on", String(event.endpoint)); - }, - }, - }, -); - -try { - // Ingestion: one borrowed sender per producer. IDs and timestamps must - // come from the source, not be regenerated on an application retry. - const events = [ - { - tradeId: "trade-12345", - timestampMs: 1723000000000, - symbol: "ETH-USD", - price: 2615.54, - amount: 0.5, - }, - { - tradeId: "trade-12346", - timestampMs: 1723000000001, - symbol: "BTC-USD", - price: 39269.98, - amount: 0.001, - }, - ]; - const sender = await db.borrowSender(); - try { - for (const event of events) { - await sender - .table("trades_sf") - .stringColumn("trade_id", event.tradeId) - .symbol("symbol", event.symbol) - .symbol("side", "buy") - .doubleColumn("price", event.price) - .doubleColumn("amount", event.amount) - .at(event.timestampMs, "ms"); - } - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence, 10_000); - } catch (error) { - if (!(error instanceof QwpIngressAckTimeoutError)) throw error; - console.warn("ACK timed out; rows remain in sf_dir for replay after close"); - } finally { - // After a terminal rejection, close() rejects with the same failure. - // Log it so that it does not replace the error thrown above. - await sender - .close() - .catch((error) => console.error("close failed:", error)); - } - - // Querying: rows may not be visible yet, see "Read-after-write". - try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query( - "SELECT timestamp, trade_id, symbol, price FROM trades_sf " + - "WHERE symbol = $1 ORDER BY timestamp DESC LIMIT 10", - { binds: (binds) => binds.setVarchar(0, "ETH-USD") }, - ); - for await (const batch of query) { - // A nonempty replay starts at sequence 0; the callback handles - // empty ones. - if (batch.batchSequence === 0n) recentPrices.length = 0; - for (const row of batch.rows()) recentPrices.push(row); - } - await query.completion; - console.log(recentPrices); - } finally { - await lease.close(); - } - } catch (error) { - if (error instanceof QwpPoolResourceError) { - // No replica was reachable within the failover budget. - console.warn("no replica available for queries:", error.cause); - } else if (error instanceof QwpEgressQueryError) { - console.error(`query failed: status=${error.status} ${error.message}`); - } else { - throw error; - } - } -} finally { - await db.close(); -} -``` - -The query can still miss newly acknowledged rows until WAL apply catches up; -use the [Read-after-write](/docs/connect/clients/nodejs/#read-after-write) pattern for a visibility guarantee. -A replayed batch is idempotent only because this example retains the event's -ID and timestamp and enables table-level deduplication. - -## Next steps - -- [Node.js client guide](/docs/connect/clients/nodejs/) for ingestion and queries. -- [Connect string reference](/docs/connect/clients/connect-string/) for the shared keys. -- [Delivery semantics](/docs/concepts/delivery-semantics/) for replay and deduplication. diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 40e37f577e..164769205d 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -2,65 +2,21 @@ slug: /connect/clients/nodejs title: Node.js client for QuestDB sidebar_label: Node.js -description: - "TypeScript and JavaScript client for QuestDB on Node.js - (@questdb/nodejs-client): QWP ingestion, streaming SQL queries, failover, and - store-and-forward." +description: "Use @questdb/nodejs-client for QWP ingestion, streaming SQL queries, failover, and store-and-forward from TypeScript or JavaScript." --- import SfDedupWarning from "../../partials/_sf-dedup-warning.partial.mdx" -The QuestDB Node.js client, `@questdb/nodejs-client`, connects Node.js -applications to QuestDB over -[QWP](/docs/connect/wire-protocols/qwp-ingress-websocket/), the QuestDB Wire -Protocol: a columnar binary protocol carried over WebSocket. The same client -ingests data at high throughput and runs SQL queries whose results stream back -as typed, column-oriented batches. - -Key capabilities: - -- **[Ingestion](#data-ingestion)**: a fluent row API and compiled, - type-checked object-row writers, with automatic table creation, schema - evolution, batching, and acknowledgement tracking. -- **[Querying](#querying)**: SQL with typed bind parameters, results streamed - as columnar batches, DDL and DML execution, cancellation, deadlines, and flow - control. -- **[One pooled client](/docs/connect/clients/nodejs-operations/#the-connection-pool)**: `connectQwpNodeClient()` - configures ingestion and queries from one `ws::` connect string, then hands - out pooled senders (`db.borrowSender()`) and query leases - (`db.borrowQuery()`). -- **[Failover](/docs/connect/clients/nodejs-operations/#failover-and-high-availability)**: multi-host endpoint lists, - automatic reconnect, and replay of unacknowledged rows. Replay is at least - once: pair it with table [deduplication](/docs/concepts/deduplication/) for - exactly-once ingestion. -- **[Store-and-forward](#store-and-forward)**: a disk journal that keeps - accepting rows while QuestDB is unreachable and survives process restarts. -- **[UDP](#fire-and-forget-udp)**: fire-and-forget ingestion for metrics where - occasional loss is acceptable. -- **[Error handling](/docs/connect/clients/nodejs-operations/#error-handling)**: typed errors, asynchronous rejection - callbacks, and connection events. The Node.js client differs from the other - QWP clients in a few places; see - [Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). - -:::tip Upgrading from 4.x or using ILP - -Version 5.0.0 adds QWP and changes how the existing `Sender` handles `null` -and `undefined` values; see [Upgrading from 4.x](/docs/connect/clients/nodejs-operations/#upgrading-from-4x). To move -existing ILP code to QWP, see [From ILP to QWP](/docs/connect/clients/nodejs-operations/#from-ilp-to-qwp). The -`Sender` class still speaks ILP over HTTP and TCP; for those transports, see -[ILP transports (legacy)](/docs/connect/clients/nodejs-operations/#ilp-transports-legacy) -on the operations and reference page. - -::: +`@questdb/nodejs-client` ingests rows and streams SQL results over the +[QuestDB Wire Protocol (QWP)](/docs/connect/wire-protocols/qwp-ingress-websocket/). +One pooled client can serve both writers and queries. It also supports the +older ILP transports for existing applications. ## Requirements -- **`@questdb/nodejs-client` 5.0.0 or newer** for QWP. Earlier versions - support ILP only. -- **Node.js 20.18.1 or newer**. -- **QuestDB 10.0.0 or newer**, which serves QWP on the HTTP port (`9000` by - default) at `/write/v4` for ingestion and `/read/v1` for queries. If QuestDB - is not running yet, see the [quick start](/docs/getting-started/quick-start/). +- `@questdb/nodejs-client` 5.0.0 or later (earlier versions support ILP only). +- Node.js 20.18.1 or later. +- QuestDB 10.0.0 or later, with QWP on its HTTP port (9000 by default). @@ -68,92 +24,56 @@ on the operations and reference page. ```shell npm install @questdb/nodejs-client@^5 -``` - -Use `yarn add @questdb/nodejs-client@^5` or -`pnpm add @questdb/nodejs-client@^5` with the other package managers. The -package exports its complete API from the package root, ships ES module and -CommonJS builds, and bundles TypeScript declarations. There are no other supported import paths. - -The examples on this page are TypeScript ES modules with top-level `await`. -To run the [quick start](#quick-start) as TypeScript, save its code as -`example.mts`, then run it from the project directory: - -```shell npm install --save-dev tsx -npx tsx example.mts ``` -The `.mts` extension enables ES modules and top-level `await` without changing -`package.json`. Run other examples the same way after supplying any required -configuration. To run them as plain JavaScript instead, use ES modules (`.mjs` -or `"type": "module"`) and remove type annotations, type-only imports, and -TypeScript assertions such as `as const`. +The examples are TypeScript ES modules: save one as `example.mts` and run +`npx tsx example.mts`. For JavaScript, use `.mjs` or `"type": "module"` +and remove TypeScript annotations. ## Quick start -Connect with one connect string, create a table, write two rows, wait until -QuestDB acknowledges them, and run a query. QuestDB applies acknowledged rows -asynchronously, so the first query may return no rows. +Create a table, publish a row, wait for QuestDB's acknowledgement, and query +it. Acknowledged rows are applied asynchronously, so the query may initially +return no rows; see [Read-after-write](#read-after-write). ```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, -} from "@questdb/nodejs-client"; +import { connectQwpNodeClient } from "@questdb/nodejs-client"; const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); try { const lease = await db.borrowQuery(); try { - // Create the table first, so the query below cannot hit a missing table. const ddl = await lease.query( "CREATE TABLE IF NOT EXISTS trades (" + - "symbol SYMBOL, side SYMBOL, price DOUBLE, amount DOUBLE, " + - "timestamp TIMESTAMP) TIMESTAMP(timestamp) PARTITION BY DAY", + "timestamp TIMESTAMP, symbol SYMBOL, side SYMBOL, " + + "price DOUBLE, amount DOUBLE" + + ") TIMESTAMP(timestamp) PARTITION BY DAY", ); await ddl.completion; - // Ingest: borrow a sender, add rows, publish them, and wait for the ACK. const sender = await db.borrowSender(); try { await sender .table("trades") .symbol("symbol", "ETH-USD") - .symbol("side", "sell") + .symbol("side", "buy") .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.00044) - .at(Date.now(), "ms"); - await sender - .table("trades") - .symbol("symbol", "BTC-USD") - .symbol("side", "sell") - .doubleColumn("price", 39269.98) - .doubleColumn("amount", 0.001) + .doubleColumn("amount", 0.5) .at(Date.now(), "ms"); await sender.flush(); await sender.waitForAcknowledged(sender.publishedSequence); } finally { - // Returns the sender to the pool. The connection stays open. - await sender.close(); + await sender.close(); // return it to the pool } - // Query. QuestDB applies acknowledged rows asynchronously, so a query - // right after ingestion can still return no rows; see Read-after-write. const query = await lease.query( - "SELECT timestamp, symbol, price, amount FROM trades " + - "WHERE symbol = 'ETH-USD' LIMIT 10", + "SELECT timestamp, symbol, price FROM trades LIMIT 10", ); for await (const batch of query) { - for (const [timestamp, symbol, price, amount] of batch.rows()) { - console.log(timestamp, symbol, price, amount); - } + for (const row of batch.rows()) console.log(row); } await query.completion; - } catch (error) { - if (!(error instanceof QwpEgressQueryError)) throw error; - // QuestDB rejected the SQL: status is the QWP status code. - console.error(`query failed: status=${error.status} ${error.message}`); } finally { await lease.close(); } @@ -162,233 +82,72 @@ try { } ``` -What happens: - -1. `connectQwpNodeClient()` validates every key of the connect string, then - opens one ingestion and one query connection. It rejects if QuestDB is - unreachable. -2. `db.borrowQuery()` leases a query connection. `lease.query()` runs one SQL - statement and returns a query handle; its `completion` promise settles when - the statement ends. -3. `db.borrowSender()` leases a sender. Rows are staged locally until an - auto-flush threshold is reached or the sender is flushed. - `waitForAcknowledged()` waits until QuestDB has committed them, and - `close()` returns the sender to the pool. -4. The query handle is an async iterable of result batches, and `batch.rows()` - yields one array per row. -5. `db.close()` closes both pools; see - [Closing the pooled client](/docs/connect/clients/nodejs-operations/#closing-the-pooled-client). - -Without the `CREATE TABLE`, the first write creates `trades` automatically, -with a designated timestamp column named `timestamp`. A `trades` table that -already exists is left unchanged, and the sender writes to its designated -timestamp. If yours uses the `trades(ts, ...)` schema from the -[PGWire guide](/docs/connect/compatibility/pgwire/nodejs/), replace -`timestamp` with `ts` in the SQL on this page. - -Timestamps come back as `bigint` microseconds since the Unix epoch; see -[Reading result values](#reading-result-values) for every type. To wait until a -write is visible to queries, see [Read-after-write](#read-after-write). +Use an event timestamp rather than `atNow()` if rows may be replayed. If a +`trades` table already exists with a different designated timestamp name, +change the SQL to match it; the +[PGWire Node.js guide](/docs/connect/compatibility/pgwire/nodejs/) uses `ts`. ## Connecting -Create a client with one of these entry points: +`connectQwpNodeClient(conf, options?)` opens a pooled ingestion and query +connection and rejects if the server is unreachable. A `ws::` connect string +uses plain WebSocket, and `wss::` uses TLS. The same `addr`, credentials, and +TLS settings apply to both directions: -| Entry point | Returns | Use it for | -|---|---|---| -| `connectQwpNodeClient(conf, options?)` | `Promise` | The recommended pooled client for ingestion and queries. Opens the pool minimums and rejects if QuestDB is unreachable. | -| `createQwpNodeClient(conf, options?)` | `QwpClient` | The same pooled client without contacting the server. It connects on `db.connect()` or on the first borrow. | -| `Sender.fromConfig(conf, options?)` | `Promise` | A standalone sender for ingestion only, or for migrating existing ILP code. | -| `connectQwpNodeSender(connection, senderOptions?, sessionOptions?)` | `Promise` | A standalone sender with every column method, configured with typed options instead of a connect string. | - -### Pooled client - -`connectQwpNodeClient()` takes one `ws::` or `wss::` connect string for both -directions. Every `addr` entry is used for ingestion (`/write/v4`) and for -queries (`/read/v1`), and the credentials and TLS keys apply to both: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;sender_pool_max=2;query_pool_max=8;", -); -try { - console.log(db.metrics.senders, db.metrics.queries); -} finally { - await db.close(); -} +```text +ws::addr=localhost:9000;sender_pool_max=2;query_pool_max=8; ``` -The `QwpClient` handle has five members: - -| Member | Returns | Purpose | -|---|---|---| -| `borrowSender()` | `Promise` | Lease an exclusive sender. Its `close()` flushes and returns it to the pool. | -| `borrowQuery()` | `Promise` | Lease an exclusive query connection. Its `close()` returns it to the pool. | -| `connect()` | `Promise` | Open the pool minimums. Called for you by `connectQwpNodeClient()`. Safe to retry after a failure. | -| `metrics` | `QwpClientMetrics` | Pool counters (`total`, `available`, `leased`, `creating`, `waiting`) for senders and queries. | -| `close()` | `Promise` | Close both pools. Resolves even if rows are not acknowledged; see [Closing the pooled client](/docs/connect/clients/nodejs-operations/#closing-the-pooled-client). Idempotent. | - -Share one `QwpClient` across your application and close it at shutdown. See -[The connection pool](/docs/connect/clients/nodejs-operations/#the-connection-pool) for pool sizing and lease rules. +`addr` accepts comma-separated or repeated hosts. A port omitted from an +address defaults to 9000. A key may appear only once (except `addr`); +unknown keys and duplicate keys fail validation. Escape a semicolon in a +value as `;;`. See the +[connect string reference](/docs/connect/clients/connect-string/) for the +shared keys and [Differences from other clients](#differences-from-other-clients) +for Node.js exceptions. ### Standalone Sender -The `Sender` class predates QWP. Changing its connect string from `http::` to -`ws::` switches it from ILP to QWP while keeping the same row API: - -```typescript -import { Sender } from "@questdb/nodejs-client"; - -const sender = await Sender.fromConfig("ws::addr=localhost:9000;"); -try { - // Opens the WebSocket now, so connection errors surface here. - await sender.connect(); - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.00044) - .at(Date.now(), "ms"); - await sender.flush(); -} finally { - // Publishes completed rows and waits up to 5 seconds for their ACK. - // Rejects with QwpSenderCloseTimeoutError if the ACK does not arrive. - await sender.close(); -} -``` - -`Sender` is ingestion-only. It accepts the complete QWP connect-string -vocabulary, and logs a warning for keys that only the pooled client can apply, -such as `query_pool_max` or `compression`. Its fluent API has only the nine -column methods that also exist for ILP, listed under -[Column methods](#column-methods). For every other QuestDB type, use its -[compiled writer](#compiled-object-row-writers) through `sender.writer()`, -which supports every type, or use a pooled sender or `connectQwpNodeSender()`, -which expose every column method. - -`connectQwpNodeSender()` builds a standalone `QwpSender` from typed options. -Its first argument takes the full ingestion URL, and credentials as an -`authorization` header value such as `` `Bearer ${token}` ``: - -```typescript -import { connectQwpNodeSender } from "@questdb/nodejs-client"; - -const sender = await connectQwpNodeSender( - { url: "ws://localhost:9000/write/v4" }, - { autoFlushRows: 5_000, autoFlushIntervalMs: 1_000 }, -); -try { - await sender - .table("orders") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .uuidColumn("order_id", "9f1c96b2-54b8-4d85-bb24-e82c6f1ac120") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); -} finally { - await sender.close(); -} -``` - -### Environment variable - -Keep credentials out of source code by putting the connect string in the -`QDB_CLIENT_CONF` environment variable: +For ingestion without a pool, use +`const sender = await Sender.fromConfig("ws::addr=localhost:9000;")`, then +`await sender.connect()` before writing and `await sender.close()` in a +`finally` block. The standalone `Sender` also supports ILP (`http::` and +`tcp::`), but exposes fewer fluent QWP column methods; use `sender.writer()` +for other types, or a pooled sender. For a typed standalone QWP sender, use +`await connectQwpNodeSender({ url: "ws://localhost:9000/write/v4" })`. -```bash -export QDB_CLIENT_CONF="wss::addr=db.example.com:9000;token=YOUR_TOKEN;" -``` +### Programmatic options -`Sender.fromEnv()` reads the variable. The pooled client takes the string -directly: +The second argument of `connectQwpNodeClient` accepts typed `sender`, +`ingressSession`, `egressSession`, `egress`, `webSocket`, `storeAndForward`, and +`pool` options. For example: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const conf = process.env.QDB_CLIENT_CONF; -if (!conf) throw new Error("QDB_CLIENT_CONF is not set"); -const db = await connectQwpNodeClient(conf); -try { - // borrow senders and query leases -} finally { - await db.close(); -} +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + sender: { awaitServerAck: true }, + ingressSession: { + onSenderError: (error) => + console.error("rejected batch:", error.category, error.serverMessage), + }, + pool: { senderPoolMax: 2 }, +}); +await db.close(); ``` -### Connect string syntax - -A QWP connect string has the form `schema::key=value;key=value;`: - -- **Schema**: `ws` (plain WebSocket) or `wss` (WebSocket over TLS). Both - default to port `9000` when `addr` omits the port. -- **`addr`**: `host[:port]`. List several endpoints for failover, either - comma-separated (`addr=a:9000,b:9000`) or by repeating the key. Enclose IPv6 - addresses in brackets: `addr=[::1]:9000`. -- **Keys** are lowercase and case-sensitive. An unrecognized key fails with - `unknown configuration key: `. Legacy ILP keys fail with a hint: - `retry_timeout` and `tls_ca` name their QWP replacements - (`reconnect_max_duration_millis` and `tls_roots`), and keys with no QWP - equivalent, such as `init_buf_size`, say that they apply only to the legacy - transports. -- **Values** end at `;`. Double a semicolon to include it in a value: - `password=p;;ssw;;rd` sets the password to `p;ssw;rd`. The trailing `;` is - optional. Every key needs a value: `client_id=;` fails with - `value is not set for 'client_id'`. -- **Each key appears once**, except `addr`. Repeating a key fails with - `Duplicate QWP cluster configuration key: ''`, and so does setting a key - and its alias, such as `user` and `username`. -- **No spaces** around the commas in `addr`: `addr=a:9000, b:9000` fails with - `Invalid QWP cluster address entry: ' b:9000'`. - -To add settings to a connect string that comes from configuration, such as -`QDB_CLIENT_CONF`, append only keys that the string does not set already, or -pass the setting as a [typed option](/docs/connect/clients/nodejs-operations/#programmatic-options), which takes -precedence without a duplicate-key error. - -The Node.js client's parser differs from some other clients in these ways: - -- `auto_flush_rows` and `auto_flush_interval` take `0`, not `off`, to disable - a trigger. `auto_flush=off` disables auto-flushing entirely. -- Size values accept the single-letter suffixes `k`, `m`, `g`, and `t` - (`sf_max_total_bytes=10g`). The two-letter forms `kb`, `mb`, and `gb` are - rejected. - -For every key and its default, see the -[connect string reference](/docs/connect/clients/connect-string/) and the -[configuration reference](/docs/connect/clients/nodejs-operations/#configuration-reference). - -## Ingestion modes {#ingestion-modes} - -The storage choice (`sf_dir`) and the first-connection choice -(`initial_connect_retry` or `lazy_connect`) are independent. This page uses -these names for how a sender publishes and retries: - -| Mode | Enabled by | `flush()` resolves when | During an outage | -|---|---|---|---| -| Default memory mode | Neither `sf_dir` nor a background start | The batch is written to the WebSocket, or queued for replay | `flush()`, auto-flushing `at()`, and a borrowed sender's `close()` wait for the reconnect, up to `reconnect_max_duration_millis` (5 minutes) | -| Background memory mode | No `sf_dir`; `initial_connect_retry=async` or `lazy_connect=on` | The batch is added to the in-memory replay queue | Rows keep being accepted until the queue is full | -| Store-and-forward | `sf_dir`, with either foreground or background startup | The batch is appended to the disk journal | Rows keep being accepted until the journal is full | - -Background startup retries the first connection indefinitely, with or without -`sf_dir`. With foreground startup, a sender with `sf_dir` must connect first; -subsequent disconnects are retried indefinitely. In the default memory mode, -a running sender stops after the reconnect budget. See -[Starting while QuestDB is down](/docs/connect/clients/nodejs-operations/#starting-while-questdb-is-down) and -[Ingestion reconnect](/docs/connect/clients/nodejs-operations/#ingestion-reconnect) for startup and outage behavior. +Typed options override the corresponding connect-string settings. A typed +`ingressSession.reconnect` or `egressSession.reconnect` object **replaces** +the entire policy from the string, not just the fields specified. For the +full typed API, see the +[client reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). ## Authentication and TLS -QWP authenticates on the WebSocket upgrade request, before any data is -exchanged. The credential and TLS keys apply to both ingestion and queries. - -### Token (Enterprise, recommended) +Use a bearer token (QuestDB Enterprise) or HTTP basic authentication. Supply +secrets through your application's configuration rather than source code: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -401,208 +160,245 @@ const db = await connectQwpNodeClient( await db.close(); ``` -The token is sent as an `Authorization: Bearer` header on every ingestion and -query upgrade. A [REST token](/docs/security/rbac/#authentication) and an OIDC -access token both use `token`. It cannot be combined with -`username`/`password`. +Basic auth uses `username=...;password=...;`. With `wss`, the client checks +certificates against Node.js's bundled CA store. For a private CA, set +`tls_roots=/path/to/ca.pem` or `NODE_EXTRA_CA_CERTS`; only PEM roots are +supported. `tls_verify=unsafe_off` disables verification for development +and cannot be combined with `tls_roots`. A token is read when the client is +created: create a new client to rotate it. + +## The connection pool + +Borrow one sender per concurrent producer and one query lease per concurrent +query. Each pool defaults to a minimum of 1 and a maximum of 4 connections; +set `sender_pool_max`, `query_pool_max`, and, if needed, their `_min` keys. +A borrowed sender's `close()` flushes completed rows and **returns it to the +pool**, but does not normally wait for their acknowledgements. A returned +lease or sender must not be reused. Size pools to the number of simultaneous +borrows, and close the client on shutdown. + +### Starting while QuestDB is down + +`connectQwpNodeClient()` normally connects at startup. `lazy_connect=on` +starts senders in the background, forces `query_pool_min=0`, and lets the +client start while QuestDB is down. A query borrowed before QuestDB is +reachable can still fail. Memory-only rows are lost if the process exits; +add `sf_dir` to keep them on disk. A locked journal fails startup even with +`lazy_connect=on`. With `sf_dir` and a background start, retries begin from +startup; without a background start, the first connection must succeed. + +### Closing the pooled client + +`db.close()` rejects new borrows, closes idle senders and queries, and waits +briefly for borrowed senders to be returned. It can resolve without every +batch being acknowledged. Wait for `sender.publishedSequence` before returning +a borrowed sender when an ACK is required, or use `sf_dir` to retain unacked +rows across restarts. In [default memory mode](#ingestion-modes), a borrowed +sender's `close()` can wait for a reconnect up to +`reconnect_max_duration_millis` (5 minutes by +default); plan your shutdown deadline accordingly. -### HTTP basic auth +## Data ingestion -```text -wss::addr=db.example.com:9000;username=admin;password=quest; -``` + + +Start a row with `table()`, add columns, and finish with +`await sender.at(timestamp, unit)` or `await sender.atNow()`. QuestDB creates +missing tables and columns automatically. The +[quick start](#quick-start) shows the full borrow/flush/close cycle. + +### Column methods + +The pooled QWP sender and `connectQwpNodeSender()` expose these methods: + +| Method | QuestDB type / value | +|---|---| +| `symbol(name, value)` | SYMBOL; use for bounded sets such as tickers and sides | +| `stringColumn(name, value)` | VARCHAR; use for high-cardinality IDs | +| `booleanColumn`, `byteColumn`, `shortColumn`, `int32Column` | BOOLEAN, BYTE, SHORT, INT | +| `longColumn`, `intColumn` | LONG; safe integer `number` or `bigint` | +| `float32Column`, `doubleColumn`, `floatColumn` | FLOAT, DOUBLE, DOUBLE | +| `timestampColumn(name, value, unit?)`, `dateColumn` | TIMESTAMP/TIMESTAMP_NS and DATE | +| `charColumn`, `binaryColumn`, `uuidColumn` | CHAR, BINARY (`Uint8Array`), UUID | +| `long256Column`, `ipv4Column`, `geohashColumn` | LONG256, IPv4, GEOHASH | +| `decimalColumnText`, `decimalColumn`, `decimal64Column`, `decimal128Column`, `decimal256Column` | DECIMAL; see [Decimals](#decimals) | +| `arrayColumn(name, value)` | Nested DOUBLE arrays; see [Arrays](#arrays) | + +`floatColumn()` writes DOUBLE and `intColumn()` writes LONG; use the `32` +variants for FLOAT and INT. A standalone `Sender` offers the nine methods +shared with ILP: `symbol`, `stringColumn`, `booleanColumn`, `floatColumn`, +`intColumn`, `timestampColumn`, `arrayColumn`, `decimalColumn`, and +`decimalColumnText`. Its compiled writer supports the other types. See the +[API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +for method signatures and accepted values. + +:::caution Use SYMBOL for bounded sets + +A sender keeps distinct SYMBOL values in a dictionary across its lifetime; +high-cardinality IDs belong in VARCHAR or UUID instead. The dictionary is +limited to 2,000,000 values. If a flush exceeds it, discard the staged rows +with `reset()` and use a new sender for new symbol values. A pooled sender is +replaced only if its `close()` fails; resetting before returning it leaves its +dictionary full. + +::: + +### Designated timestamp + +`at(value, unit)` accepts `"us"` (the default), `"ms"`, or `"ns"`. +`Date.now()` is **milliseconds**, so use `.at(Date.now(), "ms")`; +without the unit the row lands in 1970. Nanoseconds require a `bigint`. +`atNow()` asks QuestDB to assign arrival time, which changes on replay. Use +the event's timestamp for deduplication. For a newly created table, `"ns"` +creates a TIMESTAMP_NS designated timestamp; the other units create +TIMESTAMP. The default designated column name is `timestamp`. + +### Null values + +Passing `null` or `undefined` omits that column. On an existing nullable +column this stores NULL; an omitted BOOLEAN becomes `false`, and BYTE and +SHORT become `0`. An all-null column does not create a new column. Local +value errors discard the row in progress: start again with `table()`. +`cancelRow()` drops an unfinished row; `reset()` also drops rows staged since +the last flush. + + + +### Decimals + +Use `decimalColumnText(name, "0.0750")` to preserve the input scale, including +trailing zeros. Both strings and numbers accept exponents such as +`"1.5e-3"`; a JavaScript number cannot retain trailing zeros. Binary methods +`decimal64Column(name, unscaled, scale)`, `decimal128Column()`, and +`decimal256Column()` take an unscaled `bigint`. Pre-create a table if you need +a specific precision: QWP auto-creation chooses the maximum precision for the +wire width. The server currently cannot return DECIMAL with precision 9 or +less over QWP; cast it to a wider precision when querying. -`user` and `pass` are accepted aliases. Both halves must be present, and the -username cannot contain `:`. + + -### TLS +### Arrays -The `wss` schema enables TLS and, by default, verifies the server certificate -against the CA certificates bundled with Node.js, not the operating system's -trust store. A private CA installed only in the operating system is not -trusted. To trust it, -set `tls_roots`, or add it for the whole process with the -`NODE_EXTRA_CA_CERTS` environment variable, which Node.js reads at startup. -Two keys adjust verification, and both are rejected on a plain `ws` string: +`arrayColumn(name, value)` sends a uniformly shaped nested array of numbers +as DOUBLE[], DOUBLE[][], and so on. `longArrayColumn()` exists for protocol +parity, but current servers reject LONG arrays. Query results expose DOUBLE +arrays as `{ dimensions, values }`. -- `tls_roots=/path/to/ca.pem` trusts the CA certificates in a PEM file instead - of the bundled ones. The Node.js client accepts PEM only: - `tls_roots_password` and PKCS#12 or JKS stores are rejected. Export the CA - certificates to PEM first. -- `tls_verify=unsafe_off` disables certificate verification. Use it only in - development. It cannot be combined with `tls_roots`. +### Compiled object-row writers -To route the connection through an HTTP or SOCKS proxy, pass an agent such as -`https-proxy-agent` in `webSocket.agent` (or `qwp.webSocket.agent` on a -`Sender`). A custom agent owns certificate verification, so it cannot be -combined with `tls_verify` or `tls_roots`. +For a stream of objects with a fixed shape, compile a writer once with +`sender.writer(table, schema)` and call `writer.row(object)` or +`writer.rows(iterable)`. Schema builders include `symbol()`, `double()`, +`varchar()`, and `designatedTimestamp("ms")`. The writer validates each row; +its `QwpWriterRowError` names the offending table, column, and row index. +See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +for all builders. -The pooled client reports connection setup failures as the `cause` of a -`QwpPoolResourceError`. For the setup deadlines and the errors they produce, -see [Connection timeouts](/docs/connect/clients/nodejs-operations/#connection-timeouts). +### Ingestion modes {#ingestion-modes} -### Unsupported authentication paths +Storage (`sf_dir`) and the initial-connect choice are independent: -| Path | Status | Workaround | +| Mode | Enabled by | During an outage | |---|---|---| -| OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider and has no callback to refresh a token. | Obtain an access token from your identity provider, pass it as `token=...`, and create a new client before the token expires. See [OpenID Connect](/docs/security/oidc/). | -| Token rotation mid-session | Not supported. The credential is read once, when the client is created, and reused for every reconnect. QuestDB rejects an expired token when the client next opens a connection: queries and senders in default memory mode then fail, while senders with `sf_dir` or in background memory mode keep retrying and buffering (see [Connection-level errors](/docs/connect/clients/nodejs-operations/#connection-level-errors)). | Close the client and create a new one with the new token before the old one expires. | -| Mutual TLS (client certificates) | Not supported. QuestDB does not negotiate client certificates. | Use token or basic authentication over `wss`. | -| ILP JWK authentication | Not available for QWP. `auth`, `jwk`, `token_x`, and `token_y` are rejected on `ws`/`wss`. | Use token or basic authentication. | +| Default memory | Neither `sf_dir` nor a background start | `flush()` waits for reconnect, for up to 5 minutes by default | +| Background memory | No `sf_dir`; `lazy_connect=on` or `initial_connect_retry=async` | Rows queue in memory; retries continue | +| Store-and-forward | `sf_dir`, with either startup choice | Rows go to the disk journal; retries continue after the first successful connection, or from startup with a background start | -### Production example: TLS, token, and multiple hosts +### Flushing -A typical Enterprise deployment combines `wss`, a token, and several hosts in -one connect string: +Auto-flush is enabled by default: 1,000 rows or 100 ms since the last flush +(checked when a row is added, not by a background timer). Call `flush()` at +the end of a burst. `auto_flush=off` disables triggers, or set +`auto_flush_rows=0` / `auto_flush_interval=0` separately. `flush()` publishes +rows, but does not by default wait for QuestDB to accept them. To confirm +delivery, see [Awaiting acknowledgements](#awaiting-acknowledgements). -```text -wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN; -``` +#### Backpressure -Add `tls_roots=/path/to/ca.pem;` when the servers use a private CA. See -[Multiple endpoints](/docs/connect/clients/nodejs-operations/#multiple-endpoints) for routing queries to replicas, and -the [full example](/docs/connect/clients/nodejs-operations/#full-example-ingestion-and-querying-with-failover) for a -complete program with this configuration. +The replay queue defaults to 128 MiB without `sf_dir`, and the disk journal +targets 10 GiB with it. When full, publishing waits up to 30 seconds by +default, then rejects with `QwpMemoryReplayAppendTimeoutError` or +`QwpReplayStoreAppendTimeoutError`. The batch stays staged: slow down and +retry `flush()`; do not write the rows again. Configure the cap with +`sf_max_total_bytes` and the wait with `sf_append_deadline_millis`. -## Data ingestion +#### Batch size limits - +QuestDB advertises its maximum batch size on connection (about 2 MiB on a +default server). A batch too large to send fails with `QwpBatchTooLargeError`; +call `reset()` and rebuild it in smaller batches. If a sender starts offline, +it cannot yet know the server limit. With `sf_dir` and `lazy_connect=on`, set +`sf_max_segment_bytes=1m` so an oversized batch cannot block journal replay. -### General usage pattern +### Awaiting acknowledgements -A sender is not safe for concurrent producers: the row in progress is shared -state, so borrow one sender per producer (see [Concurrency](/docs/connect/clients/nodejs-operations/#concurrency)). +After `flush()`, wait for the cumulative watermark: +`await sender.waitForAcknowledged(sender.publishedSequence, 10_000)`. -1. Borrow a sender with `db.borrowSender()`, or create a - [standalone `Sender`](#standalone-sender). -2. Call `table(name)` to start a row. -3. Add values with the [column methods](#column-methods), such as - `symbol(name, value)` and `doubleColumn(name, value)`. For a nullable column, - pass `null` or `undefined`, or skip the column to store NULL (see - [Null values](#null-values) for non-nullable defaults). -4. Close the row with `at(timestamp, unit)` or `atNow()`, and `await` the - returned promise. It rejects if an auto-flush triggered by the row fails. -5. Repeat from step 2, and call `flush()` to publish staged rows. -6. `close()` the sender when done. +`publishedSequence` includes batches sent by auto-flush; `acknowledgedSequence` +is the last accepted one. `waitForAcknowledged()` rejects on timeout or server +rejection. A timeout alone does not mean the batch was rejected: it may still +be in flight. **Do not use the return value of `flushAndGetSequence()` as the +watermark for all your rows**: it returns `-1n` if an earlier auto-flush +already published them. To make each flush wait, use typed +`sender: { awaitServerAck: true }`. -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +#### Committing source offsets -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.25) - .at(Date.now(), "ms"); - await sender.flush(); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` +When consuming from Kafka or another source, record each batch's +`publishedSequence` with its last source offset. Commit only the newest offset +whose sequence is at or below `acknowledgedSequence`. Request +[durable acknowledgement](#durable-acknowledgement) if the offset must also +survive a primary failure. -`flush()` resolving means the rows were published, not that QuestDB accepted -them. QuestDB reports a rejected batch, such as one with a value of the wrong -type for an existing column, after `flush()` has resolved: to the -`onSenderError` callback, which only logs it by default, and as a rejection of -`waitForAcknowledged()`. When your code must know that QuestDB accepted the -rows, wait for the acknowledgement after flushing, with -`await sender.waitForAcknowledged(sender.publishedSequence)`. See -[Awaiting acknowledgements](#awaiting-acknowledgements) and -[Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). - -Tables and columns are created automatically, with the column types listed -below. Table and column names are validated locally with QuestDB's rules -(at most 127 UTF-8 bytes by default, see `max_name_len`), and column names are -case-insensitive: the first spelling used is kept. - -When local value validation in a column method or `at()` fails, the sender -discards the whole row in progress, including its table, so a half-built row -never reaches QuestDB. The next row must start with `table()` again; a column -method called before that throws `table name must be set before adding columns`. -`cancelRow()` discards a row in progress without an error, and `reset()` also -drops every row staged since the last flush. - -An awaited `at()` or `atNow()` can also reject because an auto-flush failed -after the row was completed. Whether the completed rows are still staged, and -what to do next, depends on the error class; see the -[Error handling](/docs/connect/clients/nodejs-operations/#error-handling) table. +### Transactions -### Column methods +Set `transaction=on` to defer server commits of auto-flushed batches until +`flush()` (or `commit()` on a pooled sender). Transactions are atomic per +table, not across tables, and QuestDB can commit early when the table exceeds +[`qwp.max.uncommitted.rows`](/docs/configuration/qwp/#qwpmaxuncommittedrows). +Closing a standalone sender without `flush()` rolls back the open transaction; +returning a pooled sender with `close()` flushes and commits instead. -These methods are available on pooled senders and on senders from -`connectQwpNodeSender()`. Each creates the listed column type when the column -does not exist yet: +### Store-and-forward -| Method | QuestDB type created | Accepted values | -|---|---|---| -| `symbol(name, value)` | SYMBOL | Any value, converted with `String()` | -| `stringColumn(name, value)` | VARCHAR | `string` | -| `booleanColumn(name, value)` | BOOLEAN | `boolean` | -| `byteColumn(name, value)` | BYTE | Integer `number` from -128 to 127 | -| `shortColumn(name, value)` | SHORT | Integer `number` from -32768 to 32767 | -| `int32Column(name, value)` | INT | 32-bit integer `number`. `-2147483648` stores NULL | -| `longColumn(name, value)`, `intColumn(name, value)` | LONG | Safe-integer `number` or `bigint`. `-9223372036854775808n` stores NULL | -| `float32Column(name, value)` | FLOAT | `number` | -| `doubleColumn(name, value)`, `floatColumn(name, value)` | DOUBLE | `number` | -| `timestampColumn(name, value, unit?)` | TIMESTAMP, or TIMESTAMP_NS with unit `"ns"` | Integer `number` or `bigint`. Unit `"us"` (default), `"ms"`, or `"ns"`; `"ns"` requires a `bigint` | -| `dateColumn(name, value)` | DATE | Epoch milliseconds as `number` or `bigint` | -| `charColumn(name, value)` | CHAR | One-character `string` (a single UTF-16 code unit) | -| `binaryColumn(name, value)` | BINARY | `Uint8Array`, copied when staged | -| `uuidColumn(name, value)` | UUID | Canonical UUID `string`, or 16 bytes in canonical big-endian order | -| `long256Column(name, w0, w1, w2, w3)` | LONG256 | Four 64-bit `bigint` words, least significant first | -| `ipv4Column(name, value)` | IPv4 | Dotted-quad `string` or packed 32-bit `number`. `0.0.0.0` is QuestDB's IPv4 NULL value and is rejected; pass `null` for NULL | -| `geohashColumn(name, bits, precisionBits)` | GEOHASH | Raw bits as `bigint`, precision from 1 to 60 bits | -| `decimalColumnText(name, value)` | DECIMAL(76, scale) | Decimal `string` or `number`. The scale comes from the literal | -| `decimalColumn(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, or big-endian two's-complement `Int8Array` | -| `decimal64Column(name, unscaled, scale)` | DECIMAL(18, scale) | Unscaled `bigint`, scale up to 18 | -| `decimal128Column(name, unscaled, scale)` | DECIMAL(38, scale) | Unscaled `bigint`, scale up to 38 | -| `decimal256Column(name, unscaled, scale)` | DECIMAL(76, scale) | Unscaled `bigint`, scale up to 76 | -| `arrayColumn(name, value)` | DOUBLE[], DOUBLE[][], ... | Nested `number` arrays of uniform shape, 1 to 32 dimensions | -| `longArrayColumn(name, value)` | LONG[] | Encoded for protocol parity. Current QuestDB servers reject LONG arrays terminally; see [Arrays](#arrays) | - -Names that differ from what you might expect: - -- `floatColumn()` and `intColumn()` write 64-bit DOUBLE and LONG. Use - `float32Column()` and `int32Column()` for FLOAT and INT. -- There is no `nullColumn()` or `setNull()`. Pass `null` or `undefined`, or - skip the column; the stored value depends on the column's - [nullability](#null-values). -- Arrays use `arrayColumn()`. `doubleArray()` is a - [compiled writer](#compiled-object-row-writers) field, not a sender method. -- `geohashColumn()` takes raw bits only. Base-32 geohash text is accepted by a - compiled writer's `geohash()` field. - -One row can mix any of these methods. This example creates an `orders` table -with a UUID, INT, LONG, BOOLEAN, and a second TIMESTAMP column: +Set `sf_dir` to journal batches across process restarts. For an offline start, +also set `lazy_connect=on` (or `initial_connect_retry=async` on a standalone +sender). Create a deduplicated table **before** ingestion if duplicates are +unacceptable: a missing table is auto-created without DEDUP. Keep event IDs +and timestamps stable across retries. + + + +```questdb-sql +CREATE TABLE IF NOT EXISTS trades_sf ( + timestamp TIMESTAMP, + trade_id VARCHAR, + symbol SYMBOL, + price DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, trade_id); +``` ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +// Persist this directory outside the container in production. +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;sf_dir=./questdb-sf;sender_id=trades;" + + "sf_max_segment_bytes=1m;lazy_connect=on;", +); try { const sender = await db.borrowSender(); try { - const submittedMs = Date.now() - 250; await sender - .table("orders") + .table("trades_sf") + .stringColumn("trade_id", "trade-12345") .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .uuidColumn("order_id", "0b2b6c4e-7c39-4f0e-9d5a-2f8e61c3a7d4") - .doubleColumn("price", 2615.54) // DOUBLE - .doubleColumn("amount", 0.5) - .int32Column("venue_id", 7) // INT - .longColumn("lots", 125n) // LONG - .booleanColumn("is_maker", true) // BOOLEAN - .timestampColumn("submitted_at", submittedMs, "ms") // TIMESTAMP - .at(Date.now(), "ms"); + .doubleColumn("price", 2615.54) + .at(1723000000000, "ms"); + await sender.flush(); // persisted locally; not necessarily acknowledged } finally { await sender.close(); } @@ -611,1674 +407,384 @@ try { } ``` -:::caution SYMBOL is for bounded sets of values - -Use SYMBOL for values from a bounded set, such as tickers, sides, or venues. -Each sender keeps every distinct SYMBOL value it has sent, across all tables -and columns, in a dictionary that holds at most 2,000,000 values. The -dictionary is not cleared while the sender lives, and no metric reports its -size. Staging a row never fails because of it: the flush that sends a batch -with a value beyond the limit fails with an `Error`, whether it is an explicit -`flush()`, the auto-flush of an `at()`, or `close()`. The rows stay staged, so -every later flush fails the same way. Call `reset()` to drop them; rows with -values the sender already knows can still be sent. Only a new sender starts -with an empty dictionary: close a standalone sender and create another. A -pooled sender whose `close()` fails with this error is replaced by the pool. -Store unique or high-cardinality values, such as trade or order IDs, as VARCHAR -with `stringColumn()`, or as UUID with `uuidColumn()`. See -[Symbol](/docs/concepts/symbol/). +`sf_durability=memory` (the default) survives a process crash, not a power +failure; `periodic` checkpoints and `append` syncs each append. The sender's +first connection is foreground by default; adding `sf_dir` alone does not +make it lazy. With `sf_dir` plus `lazy_connect=on`, it retries from startup. +A terminally rejected batch stays at the head of the journal and blocks later +rows until fixed. See the +[store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/) +and [operating guide](/docs/high-availability/store-and-forward/operating-and-tuning/). -::: +#### Journal capacity {#sf-capacity} + +`sf_max_total_bytes=10g` is a target, not a hard disk quota: transaction +completion and symbol dictionaries can exceed it. Provision extra space and +monitor the directory. Without `sf_dir`, the same key caps the memory replay +queue. + +#### Lock recovery {#sf-lock-recovery} -The standalone `Sender` class exposes only these nine: `symbol`, `stringColumn`, -`booleanColumn`, `floatColumn`, `intColumn`, `timestampColumn`, `arrayColumn`, -`decimalColumn`, and `decimalColumnText`. Its `writer()` method supports every -type. +Node.js uses a `.lock.owner` directory in each journal slot, not an OS file +lock. A crashed process can leave one behind. If opening the journal fails +with `QwpReplayStoreLockedError` (wrapped in `QwpPoolResourceError` when +pooled), verify that no other process owns the slot **before** removing a +stale lock. See the +[Node.js lock-recovery runbook](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery). +Do not let Node.js and another client's OS-lock-based sender use the same +`sf_dir` concurrently. -A column's type is fixed by the first value a sender stages for it. Writing a -different type to the same column in a later row throws -`column type mismatch for ''`. +### Durable acknowledgement -Within one row, duplicate column assignments keep the first value, including -names that differ only in case. For example, -`.doubleColumn("price", 1).stringColumn("PRICE", "wrong")` keeps `1` and does -not raise a type mismatch. Invalid values can still fail local validation. +On QuestDB Enterprise with replication, `request_durable_ack=on` makes the +acknowledgement watermark wait until the WAL has been uploaded to object +storage. `sender: { awaitDurableAck: true }` also makes each `flush()` wait. +If the server lacks support, a foreground first connection fails with +`QwpDurableAckUnavailableError` (wrapped in `QwpPoolResourceError` when +pooled). A background-started sender retries from startup and emits +`durable-ack-unavailable` connection events **even with `sf_dir`**. With +`sf_dir` and foreground startup, only later mismatches, after a successful +connection, are retried. Monitor these events and journal capacity: a +successful background start does not prove durable ACK is available. -For an existing table, QuestDB rejects an incompatible type or value -asynchronously; see [Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). Compatible -conversions are allowed: for example, `longColumn("price", 123n)` can write to -an existing DOUBLE column. This does not change the sender's local -type-consistency rule. +### Fire-and-forget UDP -### Null values +The standalone `Sender` also accepts `udp::addr=localhost:9007;` for +fire-and-forget ingestion. Enable the server's +[`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver) first. UDP has no +TLS, auth, ACK, retries, transactions, or store-and-forward; use WebSocket +for reliable writes. -Passing `null` or `undefined` to a column method omits the column, just like -leaving it out of the row. For an existing nullable column, QuestDB stores SQL -NULL. BOOLEAN, BYTE, and SHORT are not nullable: omitted BOOLEAN values become -`false`, and omitted BYTE and SHORT values become `0`. +## Querying -CHAR uses the zero character as its NULL marker. Current QWP query results can -return that marker as the one-character string `"\u0000"`, rather than JavaScript -`null`. See the [data types](/docs/query/datatypes/overview/) and -[type nullability](/docs/query/datatypes/overview/#type-nullability) references. +Borrow one query lease per concurrent query. A lease executes one query at a +time; close it in `finally`. -For example, omitting a value for an existing SYMBOL column stores NULL: +### Running a SELECT ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); try { - const sender = await db.borrowSender(); + const lease = await db.borrowQuery(); try { - const trade: { side?: string; amount?: number } = { amount: 0.011 }; - await sender - .table("trades") - .symbol("symbol", "BTC-USD") - .symbol("side", trade.side) // undefined: stored as NULL - .doubleColumn("price", 39269.98) - .doubleColumn("amount", trade.amount) - .at(Date.now(), "ms"); + const query = await lease.query( + "SELECT timestamp, symbol, price FROM trades WHERE symbol = $1 LIMIT 100", + { + binds: (binds) => binds.setVarchar(0, "ETH-USD"), + timeoutMs: 30_000, + initialCredit: 1024 * 1024, + }, + ); + for await (const batch of query) { + for (const row of batch.rows()) console.log(row); + } + await query.completion; } finally { - await sender.close(); + await lease.close(); } } finally { await db.close(); } ``` -- An omitted column is not created on a table that lacks it: a NULL carries no - type to infer from. -- The column name is still validated when the value is nullish. -- Rows that already exist in a batch, or rows added later, use the same - NULL or non-nullable default for any column they do not set. -- INT, LONG, and DATE reserve their minimum values as NULL: writing - `-2147483648` to INT or `-9223372036854775808n` to LONG or DATE stores NULL. - IPv4 reserves `0.0.0.0` for NULL too, but `ipv4Column()` rejects it with a - `RangeError` and discards the row: pass `null` to store an IPv4 NULL. -- A row where every column value is nullish is still sent over WebSocket. - Its non-designated columns use the NULL/default rules above; the designated - timestamp comes from `at()` or `atNow()`. To drop such a row instead, call - `cancelRow()` before closing it. Over UDP, `atNow()` rejects such a row while - the sender knows no non-null column for the table. +`lease.query()` returns a handle with async result batches and a `completion` +promise. For DDL/DML, await `completion` without iterating. A batch has +`rowCount`, `columns`, `rows()`, `get(rowIndex, columnIndex)`, and +`batchSequence`. A query can fail during iteration as well as at completion. -### Designated timestamp +### Reading result values + +| QuestDB type | JavaScript value | +|---|---| +| BOOLEAN, INT, DOUBLE, FLOAT | `boolean` or `number` | +| LONG | `bigint` | +| TIMESTAMP / TIMESTAMP_NS / DATE | `bigint` in microseconds / nanoseconds / milliseconds | +| VARCHAR, SYMBOL, CHAR | `string` | +| BINARY | `Uint8Array` | +| UUID | `{ low: bigint, high: bigint }` | +| DECIMAL | `{ unscaled: bigint, scale: number }` | +| DOUBLE arrays | `{ dimensions: number[], values: number[] }` | +| Nullable values | `null` (except CHAR's zero marker, which can be `"\u0000"`) | -The [designated timestamp](/docs/concepts/designated-timestamp/) controls -partitioning and ordering. Set it when closing the row: +Other types include IPv4 (signed 32-bit `number`), GEOHASH and LONG256 +objects. `JSON.stringify()` cannot serialize `bigint`: convert it to a string +first. Current servers cannot return INTERVAL, an untyped NULL, or DECIMAL +with precision 9 or less over QWP; cast those in SQL to a supported type. -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +### Bind parameters -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - // Milliseconds, for example from Date.now() - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); +Bind indexes start at 0 for `$1` and must be set in ascending order without +gaps. Use `setVarchar`, `setInt`, `setLong`, `setDouble`, +`setTimestampMicros`, `setTimestampNanos`, `setUuid`, or other typed setters +on the `binds` callback. Use `setNull(index, QWP_COLUMN_TYPE.DOUBLE)` for a +typed NULL; BINARY, IPv4, and arrays have no direct bind setter. - // Microseconds are the default unit - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "sell") - .doubleColumn("price", 2615.55) - .doubleColumn("amount", 0.2) - .at(BigInt(Date.now()) * 1000n); +### DDL and DML statements - // Server-assigned timestamp - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.56) - .doubleColumn("amount", 0.1) - .atNow(); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` +`CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` use `query()`. +For a statement without result batches, `completion.kind` is `"exec-done"`. +Only `INSERT` reliably provides a row count in `rowsAffected`; a WAL +`UPDATE` can report a transaction number instead. -`at(value, unit)` accepts an integer `number` or a `bigint` with unit `"us"` -(the default), `"ms"`, or `"ns"`. `Date.now()` returns milliseconds, so always -pass `"ms"` with it: without a unit, the value is read as microseconds and the -row lands in January 1970. Nanoseconds require a `bigint`, because epoch -nanoseconds exceed the safe integer range. When the table does not exist yet, -`"ns"` creates a `TIMESTAMP_NS` designated timestamp and the other units create -a microsecond `TIMESTAMP`. An auto-created designated timestamp column is named -`timestamp`. - -`atNow()` leaves the timestamp to QuestDB, which assigns it when the row -arrives. Rows replayed after a reconnect are stamped with the replay time. -Prefer event timestamps from your source data: they keep rows in event order and -make [deduplication](/docs/concepts/deduplication/) possible, which is -[required for exactly-once delivery](/docs/concepts/delivery-semantics/). - -Other timestamp columns use `timestampColumn(name, value, unit)` with the same -units. For converting dates and strings, see -[Date to timestamp conversion](/docs/connect/clients/date-to-timestamp-conversion/). +:::warning SQL writes can run twice -### Arrays +With `failover=on`, a lost connection can re-execute in-flight SQL, including +`INSERT`. Use `failover=off` for non-idempotent SQL and check an uncertain +outcome before retrying, or make the statement idempotent. + +::: -`arrayColumn()` takes nested `number` arrays and creates a `DOUBLE` array column -with the same number of dimensions: +### Read-after-write + +An ACK confirms commitment to the WAL, **not** query visibility: WAL apply +is asynchronous. Create the table before writing, then poll for a stable event +ID with a deadline. For the `trades_sf` table above, after publishing +`trade-12345`: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); try { - const sender = await db.borrowSender(); + const lease = await db.borrowQuery(); try { - await sender - .table("order_book") - .symbol("symbol", "BTC-USD") - // shape [2, N]: row 0 holds prices, row 1 holds sizes - .arrayColumn("bids", [ - [64901.6, 64901.5, 64901.4], - [3.02, 0.06, 1.2], - ]) - .arrayColumn("asks", [ - [64901.7, 64901.8, 64901.9], - [1.54, 0.21, 2.5], - ]) - .at(Date.now(), "ms"); + const deadline = Date.now() + 10_000; + let visible = false; + while (!visible && Date.now() < deadline) { + const query = await lease.query( + "SELECT trade_id FROM trades_sf WHERE trade_id = $1 LIMIT 1", + { + binds: (binds) => binds.setVarchar(0, "trade-12345"), + timeoutMs: Math.max(1, deadline - Date.now()), + }, + ); + for await (const batch of query) visible ||= batch.rowCount > 0; + await query.completion; + if (!visible) await new Promise((resolve) => setTimeout(resolve, 100)); + } + if (!visible) throw new Error("trade not visible in time"); } finally { - await sender.close(); + await lease.close(); } } finally { await db.close(); } ``` -Every sub-array at the same depth must have the same length, and arrays may -have 1 to 32 dimensions. Only DOUBLE arrays can be ingested: -`longArrayColumn()` exists for protocol parity, but current servers reject it -with `long arrays are not supported, only double arrays`. The rejection is -terminal; see -[Recovering from a terminal rejection](/docs/connect/clients/nodejs-operations/#recovering-from-a-terminal-rejection). -Query results return arrays as `{ dimensions, values }`; see -[Reading result values](#reading-result-values). +`failover=off` prevents a replayed query from invalidating the result mid-poll. +A fixed sleep without checking for the row is not a visibility guarantee. - +### Cancellation and timeouts -### Decimals +Set `timeoutMs` per query (or `egressSession.queryTimeoutMs` by default). +A deadline cancels the query and reports `QwpEgressQueryTimeoutError`. +Leaving a `for await` loop early also starts cancellation. `query.cancel()` +requests cancellation but does not wait for it; returning the lease waits for +the cancellation to drain up to `query_close_timeout_ms` (5 seconds by +default), then discards the connection if needed. -Create decimal columns ahead of time with the precision you need. QWP can -create them automatically, but it picks the maximum precision of the wire -width (18, 38, or 76 digits). See -[decimal data type](/docs/query/datatypes/decimal/#creating-tables-with-decimals). -To also query a decimal column over QWP, give it a precision of 10 or more: -current servers cannot return a DECIMAL with a precision of 9 or less (see -[Reading result values](#reading-result-values)). +### Flow control -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +Without a credit window, a slow consumer can buffer a large result in memory. +Set `initialCredit: 1024 * 1024` on a query, as above, or set +`initial_credit` in the connect string. The client replenishes credit as your +loop consumes batches. Use `autoCredit: false` and `query.grantCredit(bytes)` +for manual control. -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const ddl = await lease.query( - "CREATE TABLE IF NOT EXISTS trade_fees (" + - "timestamp TIMESTAMP, symbol SYMBOL, " + - "settled_price DECIMAL(18, 2), commission DECIMAL(18, 4)" + - ") TIMESTAMP(timestamp) PARTITION BY DAY", - ); - await ddl.completion; - } finally { - await lease.close(); - } +### Zero-copy result views - const sender = await db.borrowSender(); - try { - await sender - .table("trade_fees") - .symbol("symbol", "ETH-USD") - .decimal64Column("settled_price", 261554n, 2) // 2615.54 - .decimalColumnText("commission", "0.0750") // keeps the literal's scale - .at(Date.now(), "ms"); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` +For hot paths, `lease.queryViews(sql, callback)` reads typed values directly +from received bytes instead of materializing arrays. Views and byte slices +are valid only until the callback returns; copy them if you need to retain +them. See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). -- `decimalColumnText()` takes a - decimal string (such as `"0.0750"`) and preserves the literal's scale, - including trailing zeros. Strings and numbers both accept scientific notation - (such as `"1.5e-3"`); pass a string when scale matters, because JavaScript - drops trailing zeros when formatting numbers. -- - `decimalColumn(name, unscaled, scale)` takes the unscaled value as a `bigint` - or as big-endian two's-complement bytes in an `Int8Array`. -- `decimal64Column()`, `decimal128Column()`, and `decimal256Column()` take an - unscaled `bigint` and select the wire width directly. - -Scale rules: - -- The first value staged for a decimal column fixes its scale until the next - flush. Later values are rescaled exactly (`"2.50"` becomes `2.5` at scale 1), - and a value that would lose digits throws a `RangeError`, such as `"1.25"` - at scale 1. -- When QWP creates the column, the first value's scale becomes the column's - scale. -- QuestDB converts each value to the table column's scale when no digits are - lost: `"2615.5400"` is stored as `2615.54` in a `DECIMAL(18, 2)` column. A - value that would lose digits, such as `0.0015` for `DECIMAL(18, 2)`, fails - the whole batch with a terminal `schema-mismatch` rejection. Stage values - with the column's scale. +### Compression -### Compiled object-row writers +Query results default to `compression=raw`. Use `compression=zstd` (or `auto`) +for large results; `compression_level=3` is accepted only with `zstd` or +`auto`. Compression does not affect ingestion. -When your data is already a stream of objects with one shape, compile a writer -for the table once. The writer validates each complete row before staging it, -and TypeScript checks every row against the schema: +## Error handling -```typescript -import { - connectQwpNodeClient, - designatedTimestamp, - double, - QwpWriterRowError, - symbol, -} from "@questdb/nodejs-client"; +Handle the failure at the stage where it occurs: -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - const trades = sender.writer("trades", { - symbol: symbol(), - side: symbol(), - price: double(), - amount: double(), - timestamp: designatedTimestamp("ms"), - }); - - await trades.row({ - symbol: "ETH-USD", - side: "sell", - price: 2615.54, - amount: 0.00044, - timestamp: Date.now(), - }); - - // Arrays, iterables, and async iterables. - // Absent nullable fields store NULL. - await trades.rows([ - { - symbol: "BTC-USD", - side: "buy", - price: 39269.98, - timestamp: Date.now(), - }, - ]); - } catch (error) { - if (!(error instanceof QwpWriterRowError)) throw error; - // Names the table, the column, and the zero-based row index for rows(). - console.error(error.message); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` +| Failure | Action | +|---|---| +| Local value validation | Fix the value; the row in progress was discarded. Test `QwpBatchTooLargeError` before `RangeError` because it extends `RangeError`. | +| `QwpMemoryReplayAppendTimeoutError` / `QwpReplayStoreAppendTimeoutError` | The batch stays staged. Slow down and retry the flush, not the rows. | +| Server rejection | See [Ingestion errors](#ingestion-errors); a terminal rejection fails the sender. | +| `QwpEgressQueryError` | Fix SQL or bind values; the lease remains usable. | +| `QwpReconnectExhaustedError` | Close the failed sender or query lease and borrow a new one. | -A rejected row is never partly staged, and rows accepted before a failing row -in `rows()` stay staged. Unknown keys and type mismatches raise -`QwpWriterRowError`. Writers apply the sender's normal auto-flush, transaction, -and acknowledgement settings. `writer()` works on pooled senders and on the -standalone `Sender` over `ws`, `wss`, or `udp`; an ILP `Sender` throws. +### Ingestion errors -Schema fields: +A local column or `at()` validation error discards the unfinished row. +A **server** rejection can arrive after `flush()` resolves: register +`ingressSession.onSenderError` to inspect its `category`, `appliedPolicy`, +`serverStatusByte`, and `serverMessage`. Without a callback, the client logs +rejections. `waitForAcknowledged()` rejects for a rejected batch. +Branch on category, not the unstable message text, and redact messages before +sending them to external trackers: they may contain row values. Terminal +session errors also reach `ingressSession.onError` with `terminal: true`. -| Field | QuestDB type | Row value | -|---|---|---| -| `symbol()` | SYMBOL | `string` | -| `varchar()` | VARCHAR | `string` | -| `char()` | CHAR | One-character `string` | -| `bool()` | BOOLEAN | `boolean` | -| `byte()`, `short()` | BYTE, SHORT | `number` | -| `int32()` | INT | `number` | -| `int64()`, `long()` | LONG | `bigint` | -| `float32()` | FLOAT | `number` | -| `float64()`, `double()` | DOUBLE | `number` | -| `timestamp(unit)` | TIMESTAMP or TIMESTAMP_NS | `number` or `bigint`; `"ns"` requires `bigint` | -| `designatedTimestamp(unit)` | Designated TIMESTAMP, or TIMESTAMP_NS with `"ns"` | As `timestamp(unit)`, required in every row. At most one per schema. The field's key names the row property only: the value always goes to the table's designated timestamp, which is named `timestamp` when QWP creates the table | -| `date()` | DATE | Epoch milliseconds | -| `binary()` | BINARY | `Uint8Array` | -| `uuid()` | UUID | Canonical UUID `string`, 16 big-endian bytes, or `{ low, high }` | -| `long256()` | LONG256 | Unsigned 256-bit `bigint`, `0x` hex text, four little-endian words, or `{ words }` | -| `ipv4()` | IPv4 | Dotted-quad `string` or packed `number`. `0.0.0.0` is rejected; omit the field for NULL | -| `geohash(precisionBits)` | GEOHASH | Raw bits, base-32 text of `precisionBits / 5` characters, or `{ bits, precisionBits }` | -| `decimal64(scale)`, `decimal128(scale)`, `decimal256(scale)` | DECIMAL | Unscaled `bigint`, decimal text, `number`, or `{ unscaled, scale }` | -| `doubleArray()` | DOUBLE[] | Nested `number` arrays, or `{ dimensions, values }` | -| `longArray()` | LONG[] | Encoded for parity; current servers reject LONG arrays terminally, as for `longArrayColumn()` | - -LONG fields take `bigint` so they never lose precision. The object forms -(`{ low, high }`, `{ words }`, `{ bits, precisionBits }`, `{ unscaled, scale }`, -`{ dimensions, values }`) match what [query results](#reading-result-values) -return, so a queried value can be written back unchanged. A writer's decimal -field rescales values to the field's scale and rejects a value that would need -rounding. +#### Recovering from a terminal rejection -### Flushing +A terminal batch rejection stops that sender. Close it and fix the row or +schema before retrying. Without `sf_dir`, unacknowledged rows on the failed +sender are lost. With `sf_dir`, the rejected batch stays at the front of the +journal and blocks **all** tables using it until fixed or deliberately +quarantined. A pooled sender is replaced after a failed `close()`, but its +replacement sees the same blocked journal. See the +[operating guide](/docs/high-availability/store-and-forward/operating-and-tuning/). -Rows are staged in memory until a flush publishes them. Auto-flush is on by -default and flushes after the row that crosses the first threshold: +### Query errors -| Trigger | Default | Connect-string key | Typed option | -|---|---|---|---| -| Row count | 1,000 rows | `auto_flush_rows` | `autoFlushRows` | -| Time since the last flush, or since the sender was created | 100 ms | `auto_flush_interval` | `autoFlushIntervalMs` | -| Estimated buffered bytes | Disabled | `auto_flush_bytes` | `autoFlushBytes` | +SQL errors reject query iteration and `completion` with +`QwpEgressQueryError` (`status`, `message`, `requestId`). Other failures +include `QwpEgressQueryTimeoutError`, cancellation and failover exhaustion. +After cancellation, return the lease before borrowing another. Compare +`QWP_STATUS` constants instead of matching server error text. -The interval is checked when a row is added. There is no background timer, so -call `flush()` after a burst of rows, or rows staged before an idle period wait -for the next row. `auto_flush=off` disables all triggers. `auto_flush_bytes` is -clamped to 90% of the effective batch limit: the server's limit, or -`sf_max_segment_bytes` when that is lower (see -[Batch size limits](#batch-size-limits)). +### Connection-level errors -The [ingestion mode](#ingestion-modes) determines when `flush()` resolves. -In every mode, `flush()` does not wait for QuestDB to acknowledge the rows, -unless you set `awaitServerAck`. Unacknowledged batches are kept and replayed -after a reconnect. See [Awaiting acknowledgements](#awaiting-acknowledgements) -and [Store-and-forward](#store-and-forward). +The pool wraps connection creation failures in `QwpPoolResourceError`; inspect +its `cause` for an authentication `QwpUpgradeError`, a +`QwpDurableAckUnavailableError`, `QwpRoleMismatchError`, or another failure. +`QwpFailoverError.attempts` records failed endpoints. A borrow at pool +capacity times out as `QwpPoolAcquireTimeoutError`. -#### Backpressure +#### Connection timeouts -The in-memory replay queue is capped at 128 MiB. When it is full, publishing -waits up to 30 seconds for acknowledgements to free space, then rejects with -`QwpMemoryReplayAppendTimeoutError`. Tune the cap with `sf_max_total_bytes` and -the wait with `sf_append_deadline_millis`; without `sf_dir` they size the -memory queue. A store-and-forward journal applies the same backpressure and -rejects with `QwpReplayStoreAppendTimeoutError`; see -[Journal capacity](#sf-capacity). - -After an append timeout the batch stays staged and the sender stays usable. -Keep the sender, slow the producer, and call `flush()` again later. Don't write -the rows again, and don't `close()` a borrowed sender while backpressure -persists: its `close()` flushes too, so it can time out the same way, and the -pool discards a borrowed sender whose `close()` fails. In the memory modes, -every batch it held that QuestDB has not acknowledged by the time it closes is -lost with it. - -`sender.metrics.ingress` shows the backlog in every mode: `pendingReplayFrames` -and `pendingReplayBytes` count the published batches that QuestDB has not -acknowledged yet. In the memory modes, `memoryReplayUsedBytes` and -`memoryReplayMaxBytes` show how full the replay queue is, and -`totalMemoryReplayBackpressureStalls` counts publishes that had to wait. -`metrics` is available on pooled senders and on senders from -`connectQwpNodeSender()`, not on the standalone `Sender` class. +`connect_timeout` and `auth_timeout_ms` default to 15 seconds. They cover +connection setup and WebSocket upgrade; the query connection also waits for +the server's initial information frame. Set shorter timeouts if a request +needs a tighter deadline. -#### Batch size limits +## Failover and high availability -When the sender connects, QuestDB advertises the largest batch it accepts: -about 2 MiB (2,097,138 bytes) on a default server, set by -`http.recv.buffer.size`. A batch must also fit in `sf_max_segment_bytes`, -which defaults to 4 MiB with `sf_dir` and applies without `sf_dir` only when -you set it. - -A row too large to fit in one batch fails the `flush()`, or the `at()` whose -auto-flush sends it, with `QwpBatchTooLargeError` before anything is sent. -That batch can never be sent: the staged rows are kept, every later flush fails -the same way, and `close()` discards them and rejects with the same error. -Call `reset()` to drop every row staged since the last flush, then write the -rows again without the oversized one. - -Until a sender has connected once, it does not know the server's limit. This -applies in background memory mode and to a store-and-forward sender with -`lazy_connect=on` or `initial_connect_retry=async` that starts while QuestDB -is down. Batches are then capped only by -`sf_max_segment_bytes`: 4 MiB with `sf_dir`, and no cap without it. A batch -larger than the server's limit passes `flush()` but can never be delivered: -the sender keeps reconnecting, and `waitForAcknowledged()` times out. With -`sf_dir`, the batch also blocks the journal, so later rows are not delivered -and a restarted client fails with `QwpPoolResourceError`. If the client can -start while QuestDB is down, set `sf_max_segment_bytes` below the server's -limit, for example `1m`: `2m` is slightly above a default server's limit. - -### Closing a sender - -`close()` publishes the sender's completed rows and discards an unfinished row -with a warning. The rest depends on how you created the sender. With -transactions on, see also [Transactions](#transactions). - -#### Closing a standalone sender - -`close()` waits up to `close_flush_timeout_millis` (5 seconds by default) for -QuestDB to acknowledge every published row, then closes the connection. `0` or -a negative value skips the wait. If the acknowledgement does not arrive in -time, `close()` rejects with `QwpSenderCloseTimeoutError`. Its -`targetSequence` is the last sequence `close()` waited for, and its -`acknowledgedSequence` is how far QuestDB acknowledged. Without `sf_dir`, the -unacknowledged rows may be lost. With `sf_dir`, they stay in the journal for -the next sender on that directory. A rejection in `finally` replaces any error -the `try` block threw, so catch it there when that matters: +Multi-host failover requires QuestDB Enterprise replication; reconnecting to +a single restarted server also works in open source. -```typescript -import { QwpSenderCloseTimeoutError, Sender } from "@questdb/nodejs-client"; +### Multiple endpoints -const sender = await Sender.fromConfig("ws::addr=localhost:9000;"); -try { - await sender.connect(); - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.5) - .at(Date.now(), "ms"); -} finally { - try { - await sender.close(); - } catch (error) { - if (!(error instanceof QwpSenderCloseTimeoutError)) throw error; - console.warn( - `published through ${error.targetSequence}, ` + - `acknowledged through ${error.acknowledgedSequence}`, - ); - } -} -``` - -#### Closing a borrowed sender - -`close()` flushes the sender's completed rows and returns it to the pool -without closing its connection. By default it does not wait for -acknowledgements. With `awaitServerAck: true` or `awaitDurableAck: true`, the -flush performed by `close()` waits for its acknowledgement too. To confirm -delivery of every row the sender published before returning it, call -`flush()` and then `waitForAcknowledged(sender.publishedSequence)`; see -[Awaiting acknowledgements](#awaiting-acknowledgements). - -In the default memory mode, the flush in `close()` behaves like `flush()` -during an outage: it waits for the reconnect, up to -`reconnect_max_duration_millis` (5 minutes by default), then rejects with -`QwpReconnectExhaustedError`, and the rows are lost. A standalone sender's -`close()` is bounded by `close_flush_timeout_millis` instead. `db.close()` does -not wait for such a `close()` to finish, and the process stays alive until it -does. To keep shutdown within a deadline, such as a container's termination -grace period, lower `reconnect_max_duration_millis` to fit it, or use -store-and-forward: with `sf_dir`, `flush()` appends to the journal and -returns, and the rows survive the restart. - -When a borrowed sender's `close()` fails, the pool discards the sender and -opens a new one for the next borrow. In the memory modes, the batches that -QuestDB has not acknowledged by the time the discarded sender closes are lost -with it; with `sf_dir`, they stay in its journal. Because QuestDB reports rejected batches -asynchronously, a sender can fail after its `close()` already succeeded: the -error then surfaces on the next borrower's auto-flushing `at()`, `flush()`, or -`close()`, and the pool replaces the sender after that. The next borrower's own -staged rows are lost with the failed sender, even rows for other tables: write -them again on a new borrow. See [Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors). - -### Awaiting acknowledgements - -QuestDB acknowledges ingested batches asynchronously. Every published frame gets -a sequence number, and the acknowledgement watermark is cumulative, so waiting -for one sequence also covers every earlier one: - -```typescript -import { - connectQwpNodeClient, - QwpIngressAckTimeoutError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - - // Wait for every row published so far, including rows an auto-flush - // already sent. Rejects with the server's error if QuestDB rejected them. - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence, 10_000); - } catch (error) { - if (error instanceof QwpIngressAckTimeoutError) { - // Still pending in memory, but closing without sf_dir can lose them. - console.warn( - "ACK timeout; rows may be lost on close at", - error.acknowledgedSequence, - ); - } else { - throw error; - } - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -| Member | Returns | -|---|---| -| `publishedSequence` | The highest sequence this sender published, including by auto-flushes, or `-1n`. After `flush()`, it covers every row written so far. | -| `waitForAcknowledged(sequence, timeoutMs?)` | Resolves when the watermark reaches `sequence`. Rejects with `QwpIngressAckTimeoutError` on timeout (15 seconds by default), without closing the sender, or with the server's rejection. It only reads the watermark, so another task can wait while the sender keeps producing. | -| `acknowledgedSequence` | The highest acknowledged sequence, or `-1n`. It never passes a batch that QuestDB rejected. | -| `flushAndGetSequence()` | Publishes staged rows and resolves with the highest sequence (`bigint`) this call published, or `-1n` when there was nothing to publish. Rows an earlier auto-flush published are not covered. | - -:::caution Do not wait on the result of `flushAndGetSequence()` - -An auto-flush inside `at()` publishes the staged rows on its own: on the row -that reaches `auto_flush_rows`, or on the first row after the sender was idle -for `auto_flush_interval` (100 ms by default). `flushAndGetSequence()` then -has nothing left to publish and returns `-1n`, and `waitForAcknowledged(-1n)` -resolves at once, before QuestDB has acknowledged or rejected the rows. To -wait for every row written so far, call `flush()` and wait for -`publishedSequence`, as in the example above. - -::: - -To make every `flush()` wait for its acknowledgement, set `awaitServerAck`: -`connectQwpNodeClient(conf, { sender: { awaitServerAck: true } })`, or -`{ qwp: { sender: { awaitServerAck: true } } }` for a standalone `Sender`. A -server rejection then rejects the waiting `flush()` itself. See -[Ingestion errors](/docs/connect/clients/nodejs-operations/#ingestion-errors) for the error classes before and after -a terminal failure. - -Acknowledgement is not required for delivery: unacknowledged batches are -replayed after a reconnect, and a standalone sender waits for them on -`close()`. Wait for acknowledgements when your application must know that -QuestDB accepted the rows, for example before committing a source offset. If -the process exits before the acknowledgement, rows still in memory may be -lost; use [store-and-forward](#store-and-forward) to keep them across -restarts. - -#### Committing source offsets - -To commit offsets in a source such as Kafka only after QuestDB has accepted -the rows, record the sequence of each flushed batch together with the batch's -last source offset. After each flush, commit the newest offset whose sequence -is at or below `acknowledgedSequence`. The watermark is cumulative and never -passes a rejected batch, so a rejection stops further commits, and the next -`flush()` or auto-flushing `at()` reports it. The produce loop never waits for -an individual batch: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -// Stand-ins for a source such as a Kafka consumer. -async function* readSource() { - for (let offset = 0n; offset < 2_500n; offset++) { - yield { offset, price: 2615.54, amount: 0.01, timestampMs: Date.now() }; - } -} -async function commitOffset(offset: bigint) { - console.log(`committed through offset ${offset}`); -} - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const sender = await db.borrowSender(); - // Flushed batches whose last offset waits for QuestDB's acknowledgement. - const pending: { sequence: bigint; offset: bigint }[] = []; - const commitAcknowledged = async () => { - let offset: bigint | undefined; - const acknowledged = sender.acknowledgedSequence; - while (pending.length > 0 && pending[0].sequence <= acknowledged) { - offset = pending.shift()!.offset; - } - if (offset !== undefined) await commitOffset(offset); - }; - try { - let staged = 0; - let lastOffset = -1n; - const checkpoint = async () => { - await sender.flush(); - pending.push({ sequence: sender.publishedSequence, offset: lastOffset }); - staged = 0; - await commitAcknowledged(); - }; - for await (const event of readSource()) { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", event.price) - .doubleColumn("amount", event.amount) - .at(event.timestampMs, "ms"); - lastOffset = event.offset; - if (++staged === 1_000) await checkpoint(); - } - if (staged > 0) await checkpoint(); - // Before shutting down, wait for the rest and commit it. - await sender.waitForAcknowledged(sender.publishedSequence, 30_000); - await commitAcknowledged(); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -To commit as acknowledgements arrive instead of after each flush, register -`ingressSession.onProgress`. It receives `QwpIngressProgressEvent` objects whose -`kind` is `published`, `acknowledged`, or `durable-acknowledged`, with the -`sequence` they cover. Read the watermark from the event, not from the -sender: events can arrive after a borrowed sender was returned to the pool, -and a returned sender throws `QwpClientClosedError` on every access. - -### Transactions - -By default QuestDB commits each batch on its own. With transactions on, -auto-flushed batches stay in an open server-side transaction until you commit: - -```typescript -import { Sender } from "@questdb/nodejs-client"; - -const sender = await Sender.fromConfig( - "ws::addr=localhost:9000;transaction=on;auto_flush_rows=10000;", -); -try { - await sender.connect(); - for (let i = 0; i < 50_000; i++) { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", i % 2 === 0 ? "buy" : "sell") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.01) - .at(Date.now(), "ms"); - } - // Ends the transaction: QuestDB commits the auto-flushed batches and the - // staged rows together when it processes this final batch. - await sender.flush(); -} finally { - await sender.close(); -} -``` - -- The transaction is atomic per table. A flush that spans several tables - commits each table separately. -- A transaction is atomic only up to a size limit. QuestDB commits a table - early once its open transaction holds - [`qwp.max.uncommitted.rows`](/docs/configuration/qwp/#qwpmaxuncommittedrows) - rows (1,000,000 by default), and closing without `flush()` cannot roll back - what it committed. The open transaction's batches also stay in the replay - queue until the commit, so they must fit in `sf_max_total_bytes` (128 MiB - without `sf_dir`). Beyond that, publishing waits `sf_append_deadline_millis` - (30 seconds) and then rejects with `QwpMemoryReplayAppendTimeoutError`, or - `QwpReplayStoreAppendTimeoutError` with `sf_dir`. Split large loads into - several transactions. -- `flush()` ends the transaction: it publishes the final batch, and QuestDB - commits the transaction when it processes that batch. Pooled senders also - have `commit()`, an alias of `flush()`. The typed option is - `transactional: true`. -- Closing a standalone sender without calling `flush()` rolls the open - transaction back, with a warning. Tables and columns that the rolled-back - batches created remain. -- Returning a borrowed sender with `close()` commits instead, because `close()` - flushes before returning the sender to the pool. `reset()` does not prevent - this: it drops only rows staged since the last flush, not the batches already - sent in the transaction. Use a standalone sender when you may need to abandon - a transaction. -- QuestDB does not acknowledge the deferred batches until the commit, so - `waitForAcknowledged()` for a sequence inside an open transaction waits for - the commit. - -### Store-and-forward - -In the default memory mode, unacknowledged rows may be lost if the process -exits. Setting `sf_dir` turns on a disk journal instead: every batch is -appended to the journal before it is sent, a background drainer sends it in -order, and acknowledged segments are deleted. - -A frame appended to the journal but not acknowledged before a crash is sent -again, so delivery is at least once: - - - -If the table does not exist when the first batch arrives, QWP creates it -without deduplication, and replayed batches can then insert duplicate rows. -Create the table as part of a deployment or schema migration, before any -sender writes to it. A sender that starts while QuestDB is down, as in the -example below, delivers its journal as soon as QuestDB is reachable, which can -be before your own startup code gets to run DDL. If the table may already -exist without deduplication, enable it with -`ALTER TABLE ... DEDUP ENABLE UPSERT KEYS(...)`, which is safe to run again. - -Use both the event timestamp and a stable, source-assigned trade ID as upsert -keys: distinct trades can share a millisecond timestamp, symbol, and side. -Store the trade ID as VARCHAR, not SYMBOL: every trade has its own ID, and -SYMBOL is for [bounded sets of values](#column-methods). - -```questdb-sql -CREATE TABLE IF NOT EXISTS trades_sf ( - timestamp TIMESTAMP, - trade_id VARCHAR, - symbol SYMBOL, - side SYMBOL, - price DOUBLE, - amount DOUBLE -) TIMESTAMP(timestamp) PARTITION BY DAY -DEDUP UPSERT KEYS(timestamp, trade_id); -``` - -Pass the same source ID and timestamp again if the application retries an -event. The following values represent one source event; do not regenerate them -when retrying it: - -```typescript -import { - connectQwpNodeClient, - QwpPoolResourceError, - QwpReplayStoreLockedError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;" + - "sf_dir=/var/lib/my-service/qdb-sf;sender_id=ingest-a;" + - // Offline batches must fit the default server limit (about 2 MiB). - "sf_durability=append;sf_max_segment_bytes=1m;lazy_connect=on;", -).catch((error: unknown) => { - // Another process holds the journal, or a crash left a stale lock. - if ( - error instanceof QwpPoolResourceError && - error.cause instanceof QwpReplayStoreLockedError - ) { - console.error("the journal is locked; see Lock recovery"); - } - throw error; -}); -const event = { tradeId: "trade-12345", timestampMs: 1723000000000 }; -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades_sf") - .stringColumn("trade_id", event.tradeId) - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(event.timestampMs, "ms"); - // Resolves once the rows are in the journal, even if QuestDB is down. - await sender.flush(); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -With a journal, the sender keeps accepting rows while QuestDB is unreachable, -subject to [journal capacity](#sf-capacity). With `lazy_connect=on` as above, -it retries from startup; without background startup, the first connection -must succeed, then later disconnects are retried indefinitely. A new sender -opened on the same directory replays what the previous process left behind, -once it can take over the directory's lock (see [Lock recovery](#sf-lock-recovery)). - -- **Layout.** A standalone `Sender` journals into `/`. A - pooled client uses one directory per pooled sender: - `/-0`, `/-1`, and so on. `sender_id` - defaults to `default` and may contain letters, digits, `_`, and `-`. Give - every process its own `sender_id`; a second live process on the same - directory fails with `QwpReplayStoreLockedError`. A pooled client also - drains any of its own `-` journals that no pooled sender - holds, such as those left by a larger pool before a restart, without - `drain_orphans`. -- **Durability.** `sf_durability` sets how the journal reaches the disk. Its - `memory` value is unrelated to the memory ingestion modes. `memory` (the - connect-string default) relies on the operating system to write the - journal, which survives a process crash but not a power loss. `periodic` - checkpoints in the background every `sf_sync_interval_millis` (5 seconds). - `append` makes every append durable before `flush()` resolves, which adds a - disk sync to every flush: on a producer that flushes often, prefer `periodic` - or larger batches. -- **Startup.** To start the pooled client while QuestDB is down, see - [Starting while QuestDB is down](/docs/connect/clients/nodejs-operations/#starting-while-questdb-is-down). A - standalone `Sender` needs only `initial_connect_retry=async` or - `lazy_connect=on`. With the default `initial_connect_retry=off`, the first - connection must succeed. -- **Rejected batches.** A batch that QuestDB rejects terminally stays at the - head of the journal and stops ingestion through the client, for every table, - until you act; see - [Recovering from a terminal rejection](/docs/connect/clients/nodejs-operations/#recovering-from-a-terminal-rejection). -- **Orphans.** With `drain_orphans=on`, a sender also adopts and drains - journals with other `sender_id` values left under the same `sf_dir` by - processes that crashed, up to `max_background_drainers` (4) at a time. -- **Other clients.** Don't let a client in another language, such as Java, use - an `sf_dir` while a Node.js client runs on it. Those clients lock journals - with operating-system file locks, and neither kind of client sees the - other's locks, so either could open or drain a journal that the other is - writing and corrupt it. The journal format is shared: once every Node.js - client on the directory has stopped, another client can open the journals - they left behind. - -Deduplication recognizes a replayed row only when it carries the same -designated timestamp and trade ID, so reuse event values on application retries -instead of calling `atNow()` or generating a new ID. See -[Deduplication](/docs/concepts/deduplication/) for choosing keys. - -#### Journal capacity {#sf-capacity} - -With `sf_dir`, `sf_max_total_bytes` (10 GiB by default) is a journal size -target, not a hard disk limit. Transaction-closing batches can reserve extra -segments so a full journal does not block the commit needed to release space. -Segment reservations can reach roughly twice the target, depending on segment -rounding; retained symbol dictionaries and other metadata take additional -space. Provision headroom for every sender and monitor actual disk usage. -Without `sf_dir`, the key caps the in-memory replay queue instead. - -When an append cannot fit within these allowances, publishing waits up to -`sf_append_deadline_millis` (30 seconds) for acknowledgements to free space, -then rejects with `QwpReplayStoreAppendTimeoutError`; see -[Backpressure](#backpressure) for what to do next. -`sender.metrics.ingress.pendingReplayBytes` reports how much of the journal -QuestDB has not acknowledged yet. - -#### Lock recovery {#sf-lock-recovery} - -The Node.js client locks a journal directory with a `.lock.owner` directory -inside it, which records the owner's host name and process ID, instead of an -operating-system file lock. After a crash, a new sender takes over -automatically only when the owner ran on the same host and its process ID is -no longer in use. Otherwise opening the journal fails with -`QwpReplayStoreLockedError`: - -- The pooled client opens its senders' journals when it starts, so - `connectQwpNodeClient()` rejects with `QwpPoolResourceError` whose `cause` is - `QwpReplayStoreLockedError`, even with `lazy_connect=on`, and the whole - client fails to start, queries included. With `sender_pool_min=0`, the first - `borrowSender()` rejects instead. -- A standalone `Sender` rejects on `connect()`. - -This is common in containers: the application usually runs as process ID 1, -which is in use again after a restart, and a replacement container usually has -a different host name. Once you have verified that the previous owner has -exited and no process is using the slot, remove its stale -`//.lock.owner` directory and start the client again. `` is -``, or `-` for a pooled sender. If startup still -fails, a guard under `/.slot-locks` may have survived the crash too; -[Node.js lock recovery](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery) -covers it, and when this cleanup can be automated. - -For all tuning options, see -[Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/) -and the [store-and-forward keys](/docs/connect/clients/connect-string/#sf-keys). - -### Durable acknowledgement - -:::note Enterprise - -Durable acknowledgement requires QuestDB Enterprise with primary replication -configured. - -::: - -By default QuestDB acknowledges a batch when it is committed to the primary's -write-ahead log. With `request_durable_ack=on`, the acknowledgement watermark -advances only after the batch is uploaded to the replication object store, so -`waitForAcknowledged()` confirms durable upload. To make every `flush()` wait -for durability, also set the typed option `awaitDurableAck`: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const token = process.env.QDB_TOKEN; -if (!token) throw new Error("QDB_TOKEN is not set"); -const db = await connectQwpNodeClient( - `wss::addr=db.example.com:9000;token=${token};request_durable_ack=on;`, - // Every flush() waits until its batch is in the object store. - { sender: { awaitDurableAck: true } }, -); -try { - const sender = await db.borrowSender(); - try { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .doubleColumn("price", 2615.54) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); - } finally { - await sender.close(); - } -} finally { - await db.close(); -} -``` - -If the server does not support durable acknowledgement, a sender that connects -in the foreground before its first successful connection fails with -`QwpDurableAckUnavailableError`, which the pooled client reports as the -`cause` of a `QwpPoolResourceError`. A background-started sender -(`initial_connect_retry=async` or `lazy_connect=on`) instead retries from -startup and emits `durable-ack-unavailable` -[connection events](/docs/connect/clients/nodejs-operations/#connection-events), **even with `sf_dir`**. With `sf_dir` -and a foreground start, the first connection fails, but a sender that has -connected successfully before keeps retrying after a later mismatch. Monitor -these events and buffer usage: successful background startup does not confirm -that the server supports durable acknowledgement. - -### Fire-and-forget UDP - -The Node.js `Sender` can send rows as UDP datagrams, for metrics where -occasional loss is acceptable: - -```typescript -import { Sender } from "@questdb/nodejs-client"; - -const sender = await Sender.fromConfig( - "udp::addr=localhost:9007;max_datagram_size=1400;", -); -try { - await sender.connect(); - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", "buy") - .floatColumn("price", 2615.54) - .floatColumn("amount", 0.5) - .at(Date.now(), "ms"); - await sender.flush(); -} finally { - await sender.close(); -} -``` - -UDP has no authentication, TLS, acknowledgements, transactions, reconnect, or -store-and-forward. The server's UDP receiver is disabled by default; enable it -with [`qwp.udp.enabled`](/docs/configuration/qwp/#udp-receiver). The default -port is `9007`. `max_datagram_size` (1400 bytes by default) must fit your -network path. A row that cannot fit in a datagram fails the flush with -`QwpUdpDatagramTooLargeError`. As with an -[oversized WebSocket batch](#flushing), the staged rows are kept, so later -flushes and `close()` fail too: call `reset()` to drop them. -`multicast_ttl` sets the multicast time-to-live. - -## Querying - -Queries run on a lease borrowed from the pooled client. One lease runs one -query at a time, so borrow one lease per concurrent query. - -### Running a SELECT - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query( - "SELECT timestamp, symbol, price, amount FROM trades " + - "WHERE symbol = $1 AND price > $2 LIMIT 100", - { - binds: (binds) => binds.setVarchar(0, "ETH-USD").setDouble(1, 2000), - timeoutMs: 30_000, - }, - ); - for await (const batch of query) { - for (const [timestamp, symbol, price, amount] of batch.rows()) { - console.log(timestamp, symbol, price, amount); - } - } - const completion = await query.completion; - if (completion.kind === "result-end") { - console.log("rows:", completion.totalRows); - } - } catch (error) { - if (!(error instanceof QwpEgressQueryError)) throw error; - console.error(`query failed: status=${error.status} ${error.message}`); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} +```text +wss::addr=db-a.example.com:9000,db-b.example.com:9000;token=YOUR_TOKEN; ``` -`lease.query(sql, options?)` sends the query and resolves with a -`QwpEgressQuery` handle. The options are a `QwpEgressQueryOptions` object: - -| Query option | Default | Purpose | -|---|---|---| -| `binds` | none | Callback that sets the `$1`, `$2`, ... parameters. See [Bind parameters](#bind-parameters). | -| `timeoutMs` | session `queryTimeoutMs` (none) | Deadline that cancels the query. It covers the whole query, including a re-execution after failover. `0` disables it. | -| `initialCredit` | session value (`0`, unbounded) | Flow-control window in bytes. Without one, the client holds whatever the server sends ahead of your loop in memory, so set it for large results. See [Flow control](#flow-control). | -| `autoCredit` | `true` | Replenish the credit window as batches are consumed. | -| `resetDictionary` | `false` | Ask the server to reset its symbol dictionary for this connection first. | - -There is no per-query failover setting; see [Query failover](/docs/connect/clients/nodejs-operations/#query-failover). -The `QwpEgressQuery` handle has these members: - -| Member | Purpose | -|---|---| -| `for await (const batch of query)` | Yields `QwpResultBatch` objects in order. | -| `completion` | `Promise` that settles when the query ends. | -| `cancel()` | Asks QuestDB to stop the query; see [Cancellation and timeouts](#cancellation-and-timeouts). | -| `awaitCompletion(timeoutMs)` | Resolves `false` if the query is still running after `timeoutMs`. | -| `isDone()` | Whether the query has ended. | -| `grantCredit(bytes)` | Adds flow-control credit when `autoCredit` is `false`. | -| `requestId` | The `bigint` that numbers queries on this connection. | - -Iteration and `completion` reject with the same error when the query fails. -Consume the result through `for await`, or `await query.completion` directly -for statements that return no rows. - -A `QwpResultBatch` has: - -- `rowCount` and `columns`: an array of `{ name, type, values, scale?, precisionBits? }`, - where `values` holds one entry per row and `type` is the numeric QWP type - code. Compare it with the exported `QWP_COLUMN_TYPE` constants: `BOOLEAN`, - `BYTE`, `SHORT`, `CHAR`, `INT`, `LONG`, `FLOAT`, `DOUBLE`, `SYMBOL`, - `VARCHAR`, `TIMESTAMP`, `TIMESTAMP_NANOS`, `DATE`, `UUID`, `LONG256`, - `GEOHASH`, `IPV4`, `BINARY`, `DOUBLE_ARRAY`, `LONG_ARRAY`, `DECIMAL64`, - `DECIMAL128`, and `DECIMAL256`. -- `rows()`: a generator that yields one array of values per row. -- `get(rowIndex, columnIndex)`: one value. -- `batchSequence`: the batch's position in the result, starting at `0n`. - -Batch objects stay valid after iteration moves on, so you can keep them. - -With the default `failover=on`, a lost connection can make the client run the -query again from its first batch. If your loop accumulates rows, reset them -when `batch.batchSequence === 0n`. A replay that returns **no batches** has no -sequence to detect, so also clear accumulated state on `onReplayReset` (on a -client with only one active query), or use `failover=off` and retry the whole -query. See [Query failover](/docs/connect/clients/nodejs-operations/#query-failover). +Ingestion needs the primary; queries can use any healthy node. `target` +filters the roles queries accept (`any`, `primary`, `replica`), while `zone` +prefers same-zone nodes. On Node.js, setting `target=replica` **in the shared +connect string also filters ingestion**, so use the typed query-only option +`{ egress: { target: "replica" } }` instead. `target=replica` is strict, not +"prefer replica and fall back to primary". If no replica is up at startup, +set `query_pool_min=0` to defer the query connection. + +### Ingestion reconnect + +Senders resend unacknowledged batches after a disconnect. A sender in default +memory mode gives up after `reconnect_max_duration_millis` (5 minutes by +default); background memory mode or `sf_dir` retries indefinitely, subject +to queue or journal capacity. The first connection is fail-fast unless +`initial_connect_retry=on` (bounded), `initial_connect_retry=async`, or +`lazy_connect=on` (background). Setting a `reconnect_*` key implicitly +requests bounded first-connection retry unless you explicitly set +`initial_connect_retry=off`. + +### Query failover + +A lost query connection can re-execute the query **from its first row**, even +if your loop has already consumed rows. A replay can also return zero batches. +For streaming results you cannot retract, use `failover=off` and retry the +whole operation after a transport failure. Otherwise buffer the result and +reset it using `egressSession.onReplayReset`; use a single-query client for +that callback because request IDs are per connection, not unique across +pooled leases. `batch.batchSequence === 0n` detects a nonempty replay but +not a replay returning no batches. Query failover defaults to 8 attempts, +which may end before the 30-second time budget: raise +`failover_max_attempts` for longer outages. A failed query lease must be +closed and replaced. + +### Typed reconnect policy + +Typed `ingressSession.reconnect` and `egressSession.reconnect` **replace** +their respective connect-string policies; set every limit you rely on in +the typed object. Ingestion `reconnect_*` keys trigger first-connection +retry, but a typed ingestion `reconnect` object does not. A query's first +connection retries only if `failover=on` is explicit, a `failover_*` key is +set without `failover=off`, or a typed query `reconnect` object is provided. + +### Connection events + +Set `ingressSession.reconnect.onEvent` or +`egressSession.reconnect.onEvent` to observe `connected`, `reconnecting`, +`reconnected`, `failed-over`, and `durable-ack-unavailable` events. No event +marks a terminal failure: use `ingressSession.onError` with `terminal: true` +for that. Connect-string `connection_listener_inbox_capacity` configures +ingestion events only; for query events, use typed +`egressSession.connectionListenerInboxCapacity`. + +## Concurrency + +Share one `QwpClient`, but keep one sender per producer and one query lease per +concurrent query. Worker threads need their own clients and, with `sf_dir`, +distinct `sender_id` values. -### Reading result values - -Values arrive as these JavaScript types: - -| QuestDB type | JavaScript value | -|---|---| -| BOOLEAN | `boolean` | -| BYTE, SHORT, INT | `number` | -| FLOAT, DOUBLE | `number` | -| LONG | `bigint` | -| TIMESTAMP | `bigint` microseconds since the Unix epoch | -| TIMESTAMP_NS | `bigint` nanoseconds since the Unix epoch | -| DATE | `bigint` milliseconds since the Unix epoch | -| CHAR | one-character `string` | -| VARCHAR, STRING, SYMBOL | `string` | -| BINARY | `Uint8Array` | -| IPv4 | `number`, as a signed 32-bit integer: `192.168.0.1` arrives as `-1062731775`. Use `value >>> 0` for the unsigned address | -| UUID | `{ low: bigint, high: bigint }`, the unsigned low and high 64-bit halves | -| LONG256 | `{ words: [bigint, bigint, bigint, bigint] }`, least significant word first. Each word is a signed 64-bit value: `BigInt.asUintN(64, word)` gives its unsigned value | -| GEOHASH | `{ bits: bigint, precisionBits: number }` | -| DECIMAL with a precision of 10 or more | `{ unscaled: bigint, scale: number }`: the value is `unscaled / 10^scale` | -| DOUBLE[], DOUBLE[][], ... | `{ dimensions: number[], values: number[] }` with values in row-major order | -| NULL in nullable types other than CHAR | `null` | - -BOOLEAN, BYTE, and SHORT are non-nullable, so omitted values read back as -`false`, `0`, and `0`. A CHAR NULL marker can currently read back as the -one-character string `"\u0000"`, not JavaScript `null`. See -[Null values](#null-values). - -Some column types cannot be returned over QWP. The server rejects such a query -with status `0x06` and a message such as `unsupported column type INTERVAL`. -Convert the column in SQL instead: - -- INTERVAL: select the bounds with `interval_start()` and `interval_end()`, - which return timestamps, or cast the interval with `::varchar`. -- DECIMAL with a precision of 9 or less, which QuestDB stores as DECIMAL8, - DECIMAL16, or DECIMAL32: cast it to a wider precision, for example - `price::DECIMAL(18, 2)`. -- An untyped `NULL` literal, as in `SELECT NULL`: give it a type, for example - `NULL::double`. - -`JSON.stringify()` throws on `bigint`, which LONG, TIMESTAMP, DATE, UUID, -LONG256, and DECIMAL values contain, so convert rows before serializing them, -as `toJson()` does below. Converting common types: - -```typescript -// TIMESTAMP (bigint microseconds) to Date. Drops sub-millisecond precision. -const toDate = (micros: bigint) => new Date(Number(micros / 1000n)); - -// UUID to its canonical string form -function uuidToString({ low, high }: { low: bigint; high: bigint }): string { - const hex = - high.toString(16).padStart(16, "0") + low.toString(16).padStart(16, "0"); - return [ - hex.slice(0, 8), - hex.slice(8, 12), - hex.slice(12, 16), - hex.slice(16, 20), - hex.slice(20), - ].join("-"); -} - -// IPv4 (signed number) to dotted quad -const ipv4ToString = (value: number) => - [24, 16, 8, 0].map((shift) => ((value >>> 0) >>> shift) & 0xff).join("."); - -// Any row or value to JSON, with bigint values as decimal strings -const toJson = (value: unknown) => - JSON.stringify(value, (_key, v) => - typeof v === "bigint" ? v.toString() : v, - ); - -console.log(toDate(1723000000000000n).toISOString()); -console.log( - uuidToString({ low: 13485158461794337056n, high: 11465204444048149893n }), -); -console.log(ipv4ToString(-1062731775)); -console.log(toJson([1723000000000000n, "ETH-USD", 2615.54])); -``` - -### Bind parameters + -Bind values are set by a callback on a `QwpBindValues` object. Indexes are -zero-based (index `0` is `$1`), and setters must be called in ascending index -order without gaps. Every setter returns the object, so calls chain: +## Configuration reference -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +The [connect string reference](/docs/connect/clients/connect-string/) lists +shared keys and defaults. Important Node.js differences: -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query( - "SELECT timestamp, symbol, price FROM trades " + - "WHERE symbol = $1 AND side = $2 AND timestamp >= $3 LIMIT $4", - { - binds: (binds) => - binds - .setVarchar(0, "ETH-USD") - .setVarchar(1, "buy") - .setTimestampMicros(2, BigInt(Date.now() - 3_600_000) * 1000n) - .setLong(3, 1000), - }, - ); - for await (const batch of query) { - for (const row of batch.rows()) console.log(row); - } - await query.completion; - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` +### Differences from other clients -| Setter | Bind type | +| Area | Node.js behavior | |---|---| -| `setBoolean(index, value)` | BOOLEAN | -| `setByte(index, value)` | BYTE | -| `setShort(index, value)` | SHORT | -| `setChar(index, value)` | CHAR (one-character `string`) | -| `setInt(index, value)` | INT | -| `setLong(index, value)` | LONG (`number` or `bigint`) | -| `setFloat(index, value)` | FLOAT | -| `setDouble(index, value)` | DOUBLE | -| `setDate(index, millis)` | DATE (`number` or `bigint`) | -| `setTimestampMicros(index, micros)` | TIMESTAMP (`number` or `bigint`) | -| `setTimestampNanos(index, nanos)` | TIMESTAMP_NS (`number` or `bigint`) | -| `setVarchar(index, value)` | VARCHAR, STRING, and SYMBOL comparisons. `null` binds NULL | -| `setUuid(index, value)` or `setUuid(index, low, high)` | UUID, as a canonical string or two 64-bit halves. `null` binds NULL | -| `setLong256(index, w0, w1, w2, w3)` | LONG256, least significant word first. Each word is a `number` or `bigint` | -| `setGeohash(index, precisionBits, value)` | GEOHASH. `value` is a `number` or `bigint` | -| `setDecimal64(index, scale, unscaled)` | DECIMAL64. `unscaled` is a `number` or `bigint` | -| `setDecimal128(index, scale, low, high)` | DECIMAL128. Each half is a `number` or `bigint` | -| `setDecimal256(index, scale, w0, w1, w2, w3)` | DECIMAL256. Each word is a `number` or `bigint` | -| `setNull(index, type)` | A typed NULL. `type` is a `QWP_COLUMN_TYPE` constant for a scalar type, for example `setNull(1, QWP_COLUMN_TYPE.DOUBLE)`. The TIMESTAMP_NS constant is `TIMESTAMP_NANOS`. BINARY, IPv4, arrays, and SYMBOL are excluded (bind text as VARCHAR). | -| `setNullDecimal64/128/256(index, scale)`, `setNullGeohash(index, precisionBits)` | NULL decimals and geohashes, which carry a scale or precision | - -There is no setter for BINARY, IPv4, or arrays. Bind IPv4 as a string and cast -it in SQL (`WHERE ip = $1::ipv4` with `setVarchar`), and pass array values as -SQL literals. - -Decimal and geohash setters take the scale or precision before the value, the -reverse of the matching column methods: `setDecimal64(index, scale, unscaled)` -but `decimal64Column(name, unscaled, scale)`, and -`setGeohash(index, precisionBits, value)` but -`geohashColumn(name, bits, precisionBits)`. - -### DDL and DML statements - -`CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` go through the -same `query()` call. They produce no batches, and `completion` resolves with -`kind: "exec-done"` instead of `kind: "result-end"`. - -:::warning DDL and DML can run twice with query failover - -With the default `failover=on`, a connection loss replays any in-flight SQL, -including DDL and DML. QuestDB may have applied an `INSERT` before its -`exec-done` response was lost, so replay can insert it again. A transport -error does not prove the statement failed. Use a separate client with -`failover=off` for non-idempotent statements, as below, and verify an -uncertain outcome before retrying manually. Alternatively, make the statement -idempotent so that a second run changes nothing, for example -`CREATE TABLE IF NOT EXISTS`, or an `INSERT` of rows with stable key values -into a table with [deduplication](/docs/concepts/deduplication/). - -::: - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); -try { - const lease = await db.borrowQuery(); - try { - const statements = [ - "CREATE TABLE IF NOT EXISTS fills (" + - "timestamp TIMESTAMP, symbol SYMBOL, side SYMBOL, " + - "price DOUBLE, amount DOUBLE" + - ") TIMESTAMP(timestamp) PARTITION BY DAY", - "INSERT INTO fills VALUES (now(), 'ETH-USD', 'buy', 2615.54, 0.5)", - "UPDATE fills SET amount = 0.6 WHERE symbol = 'ETH-USD'", - ]; - for (const sql of statements) { - const statement = await lease.query(sql); - const completion = await statement.completion; - // rowsAffected counts rows only for INSERT; see the table below. - if (completion.kind === "exec-done" && sql.startsWith("INSERT")) { - console.log(`inserted ${completion.rowsAffected} rows`); - } - } - } catch (error) { - if (!(error instanceof QwpEgressQueryError)) throw error; - console.error(`statement failed: status=${error.status} ${error.message}`); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -| `completion.kind` | Returned for | Fields | -|---|---|---| -| `"result-end"` | Queries that return rows | `totalRows` (`bigint`) | -| `"exec-done"` | DDL and DML | `rowsAffected` (`bigint`, rows written by an `INSERT`), `operationType` (QuestDB's numeric statement type) | - -Only `INSERT` reliably reports a row count in `rowsAffected`. For an `UPDATE` -on a WAL table, the default, it currently holds a transaction number rather -than the number of rows changed, so do not use it to check whether an `UPDATE` -matched any rows. On a non-WAL table, `UPDATE` reports the rows changed. DDL -reports `0`, except `TRUNCATE`, which currently reports -`18446744073709551615n`. - -Statements run in order on one lease, because each is awaited before the next -starts, so a `CREATE TABLE` is complete before the `INSERT` that follows it. - -### Read-after-write - -When `flush()` resolves, the client has published the rows, but QuestDB may not -have received them yet. QuestDB acknowledges a batch once it has committed it -to its write-ahead log, and applies committed rows to the table -asynchronously. A query that runs right after ingestion can therefore fail -with `table does not exist` on a first run, or succeed and return no rows. - -When your code must read its own writes, create the table first, write an event -with a unique ID, and poll for **that ID**. Pre-creating the table avoids -mistaking an unrelated SQL error for the first-write table-creation delay. Give -each query the time remaining until the deadline so a stalled query cannot -leave the poll running indefinitely: - -```typescript -import { randomUUID } from "node:crypto"; -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -let visible = false; -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;query_pool_max=1;", - // There is only one active query. Clear its state even if replay returns no rows. - { - egressSession: { - onReplayReset: () => { - visible = false; - }, - }, - }, -); -try { - const lease = await db.borrowQuery(); - try { - const ddl = await lease.query( - "CREATE TABLE IF NOT EXISTS trades_readback (" + - "timestamp TIMESTAMP, trade_id VARCHAR, symbol SYMBOL" + - ") TIMESTAMP(timestamp) PARTITION BY DAY", - ); - await ddl.completion; - - const tradeId = randomUUID(); - const sender = await db.borrowSender(); - try { - await sender - .table("trades_readback") - .stringColumn("trade_id", tradeId) - .symbol("symbol", "ETH-USD") - .at(Date.now(), "ms"); - await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence); - } finally { - await sender.close(); - } - - const deadline = Date.now() + 10_000; - while (!visible) { - const remainingMs = deadline - Date.now(); - if (remainingMs <= 0) throw new Error("trade not visible in time"); - const query = await lease.query( - "SELECT trade_id FROM trades_readback WHERE trade_id = $1 LIMIT 1", - { - binds: (binds) => binds.setVarchar(0, tradeId), - timeoutMs: remainingMs, - }, - ); - for await (const batch of query) visible ||= batch.rowCount > 0; - await query.completion; - if (!visible) { - await new Promise((resolve) => - setTimeout( - resolve, - Math.min(100, Math.max(0, deadline - Date.now())), - ), - ); - } - } - console.log(`visible trade: ${tradeId}`); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -Because the table already exists, an SQL error fails the loop at once instead -of being retried; the loop retries only while the row is not yet visible. The -single-query pool lets `onReplayReset` reset `visible` even when a replay -returns no batches. For concurrent queries, use separate clients for this -pattern, or set `failover=off` and retry the entire poll after a connection -error. Do not replace the poll with a fixed sleep: apply latency varies with -load. - -### Cancellation and timeouts - -A query ends early in four ways: - -- **Deadline.** Set a default with `egressSession: { queryTimeoutMs }`, or per - query with `timeoutMs`. On expiry, iteration and `completion` reject with - `QwpEgressQueryTimeoutError`, and the client cancels the query in the - background. -- **Leaving the loop.** `break`, `return`, or an exception inside `for await` - cancels the query in the background, and `completion` rejects with - `QwpEgressQueryAbandonedError`. Use it to stop reading a result at once. -- **Cancel.** `await query.cancel()` asks QuestDB to stop and returns without - waiting for it. Keep consuming the result afterwards: QuestDB acts on the - cancel only while the result is moving, and iteration and `completion` then - reject with `QwpEgressQueryError` whose `status` is `0x0a` - (`QWP_STATUS.CANCELLED`). If you stop consuming and only await `completion`, - it rejects with `QwpEgressQueryCancelTimeoutError` after - `query_close_timeout_ms` (5 seconds), and the client closes the connection. - A query that finishes before QuestDB processes the cancel, or a DDL or DML - statement, completes normally instead. -- **Waiting without cancelling.** `await query.awaitCompletion(timeoutMs)` - resolves `false` when the wait times out and leaves the query running. - `query.isDone()` reports whether the query has ended. - -QuestDB checks for a cancel between result batches, while it is sending. Set a -[credit window](#flow-control), with `initial_credit` in the connect string or -`initialCredit` per query, so that QuestDB pauses when your loop falls behind -and stops at the next batch after a cancel, a deadline, or an early exit. -Start with 1 MiB; the client replenishes it as your loop consumes batches. -Without a credit window, QuestDB streams as fast as the network allows: it may -send the whole result before it reads a cancel, and the client holds whatever -arrived. A cancelled query can then complete normally, and returning the lease -after an early exit waits while the rest of the result streams. If draining -that backlog takes longer than `query_close_timeout_ms`, the cancel fails with -`QwpEgressQueryCancelTimeoutError` and the connection is closed, even while -your loop keeps consuming. A credit window does not interrupt expensive work -before the next batch is ready. - -```typescript -import { - connectQwpNodeClient, - QwpEgressQueryTimeoutError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query( - "SELECT symbol, avg(price) FROM trades SAMPLE BY 1m", - // With credit, QuestDB stops at the next batch after the deadline. - { timeoutMs: 5_000, initialCredit: 1024 * 1024 }, - ); - for await (const batch of query) { - console.log(batch.rowCount); - } - } catch (error) { - if (!(error instanceof QwpEgressQueryTimeoutError)) throw error; - console.warn( - `query ${error.requestId} timed out after ${error.timeoutMs} ms`, - ); - } finally { - // Waits for the cancellation to drain before the lease is reused. - await lease.close(); - } -} finally { - await db.close(); -} -``` - -To stop a result after enough rows, call `cancel()` and keep iterating until -the result ends: - -```typescript -import { - connectQwpNodeClient, - QWP_STATUS, - QwpEgressQueryError, -} from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM trades", { - initialCredit: 1024 * 1024, - }); - let rows = 0; - let cancelled = false; - try { - for await (const batch of query) { - rows += batch.rowCount; - if (rows >= 10_000 && !cancelled) { - cancelled = true; - // Keep iterating: the result ends a few batches later. - await query.cancel(); - } - } - await query.completion; - } catch (error) { - const wasCancelled = - error instanceof QwpEgressQueryError && - error.status === QWP_STATUS.CANCELLED; - if (!wasCancelled) throw error; - } - console.log(`read ${rows} rows`); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -After a query ends early, its connection stays busy until QuestDB confirms the -cancellation, and another `query()` on the same lease rejects with -`a QWP query is already active on this connection`. Return the lease with -`close()` and borrow a new one for the next query. `close()` waits up to -`query_close_timeout_ms` (5 seconds) for QuestDB to confirm the cancellation. -If it does not, `close()` discards the connection, which can take as long -again, so returning the lease can take up to twice `query_close_timeout_ms`. - -### Flow control - -By default QuestDB streams results as fast as the network allows. The client -decodes up to four batches ahead of your loop (`buffer_pool_size`), but it keeps -every frame it receives in memory until your loop consumes it, so a slow loop -over a large result can hold most of that result in memory. To bound how much -the server sends ahead, set a byte-credit window with `initial_credit` in the -connect string or `initialCredit` per query. Start with 1 MiB for large or -unbounded results: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); -try { - const lease = await db.borrowQuery(); - try { - const query = await lease.query("SELECT * FROM trades", { - initialCredit: 1024 * 1024, // server pauses after about 1 MiB - }); - for await (const batch of query) { - // The client replenishes the credit as each batch is consumed. - console.log(batch.rowCount); - } - await query.completion; - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -With `autoCredit: false`, call `query.grantCredit(bytes)` yourself. To cap the -rows in each batch, set `max_batch_rows` (1 to 1,048,576), or the typed option -`egress: { maxBatchRows }`; it applies to every query. A credit window -also lets QuestDB stop at the next batch after a cancel, a deadline, or an -early exit; see [Cancellation and timeouts](#cancellation-and-timeouts) for -what your loop must do after `cancel()`. - -### Zero-copy result views - -`query()` materializes every value into JavaScript arrays. For hot paths, -`queryViews()` hands a reusable view of each batch to a callback, and reads -values straight from the received bytes. This single-pass sum disables query -failover so a transport error rejects instead of leaving a partial sum: - -```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; - -const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); -try { - const lease = await db.borrowQuery(); - try { - let notional = 0; - const query = await lease.queryViews( - "SELECT timestamp, symbol, price, amount FROM trades", - (batch) => { - const price = batch.column(2); - const amount = batch.column(3); - for (let row = 0; row < batch.rowCount; row++) { - if (!price.isNull(row) && !amount.isNull(row)) { - notional += price.getDouble(row) * amount.getDouble(row); - } - } - }, - ); - await query.completion; - console.log({ notional }); - } finally { - await lease.close(); - } -} finally { - await db.close(); -} -``` - -For row-major work instead, `batch.forEachRow()` reuses one row object; read -values inside its callback only when they contribute to your result. For an -accumulator that supports automatic query replay, see -[Query failover](/docs/connect/clients/nodejs-operations/#query-failover). - -Batches are delivered one at a time: when the callback returns a promise, the -client waits for it before delivering the next batch. With a credit window set -(`initialCredit`), it also grants credit for a batch only after its callback -resolves, so a slow callback throttles the server. - -The batch, its column views, and any `Uint8Array` returned from them are valid -only until the callback returns: copy a byte view with `.slice()`, or call -`batch.materialize()`, to keep data. Column views provide typed getters such as -`getBoolean`, `getInt`, `getLong`, `getDouble`, `getString`, `getSymbol`, -`getBinaryView`, and `get` for any type. - -### Compression - -Ask for zstd-compressed results to save bandwidth on large result sets: - -```text -ws::addr=localhost:9000;compression=zstd;compression_level=3; -``` - -`compression` is `raw` (the default), `zstd`, or `auto`; `zstd` and `auto` both -accept a raw reply. `compression_level` ranges from 1 to 22, and the server may -clamp it. `lease.negotiatedCompression` reports what the server chose, for -example `{ codec: "zstd", level: 3 }`. Compression applies to query results -only. - -### Server information - -`lease.serverInfo` describes the server the lease is connected to: `role` (a -`QWP_SERVER_ROLE` value: standalone, primary, replica, or primary catching up), -`zoneId`, `clusterId`, `nodeId`, and `capabilities`. It refreshes after a -failover. - - - -For typed options and the key table, see the -[Node.js configuration reference](/docs/connect/clients/nodejs-operations/#configuration-reference). +| Startup/outage | `lazy_connect=on` starts senders in the background; default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. | +| Query startup | First-connect retry requires explicit `failover=on`, a `failover_*` key (without `failover=off`), or typed `egressSession.reconnect`. | +| `target`, `zone` | Also apply to ingestion when set in the connect string; use typed `egress.target` for queries only. | +| `sf_dir` | Node.js recursively creates missing parents and a slot. Its `.lock.owner` directory can outlive a crashed process; see [Lock recovery](#sf-lock-recovery). | +| SF-only keys | Explicit `sf_durability` (even `memory`), `sf_sync_interval_millis`, `drain_orphans`, `max_background_drainers`, and `catch_up_cap_gap_min_escalation_window_millis` require `sf_dir`. `sf_durability=append` is supported. | +| `sf_max_total_bytes` | With `sf_dir` it is a journal size **target**, not a disk quota; without `sf_dir` it caps the memory queue. | +| Durable ACK | Background-started senders, including with `sf_dir`, retry an unavailable capability from startup. Explicit `durable_ack_keepalive_interval_millis` also requests durable ACK even at `0`; negatives are rejected. | +| `connection_listener_inbox_capacity` | Configures only ingestion events, not query events. | +| Parsing | Use `0`, not `off`, for `auto_flush_rows` and `auto_flush_interval`; sizes take single-letter suffixes (`k`, `m`, `g`, `t`). `compression_level` requires `compression=zstd` or `auto`. `tls_roots` must be PEM; `tls_roots_password` is rejected. | +| Reconnect jitter | Ingestion uses full jitter, unlike the shared guide's equal-jitter ingestion schedule. | +| Typed SF options | A typed-only `storeAndForward` object defaults to append durability, 1 GiB, and fail-on-full; connect strings default to memory durability, 10 GiB, and wait-on-full. | +| Standalone `Sender` | Ignores pool and query-only keys with a warning; `on_*_error` keys are accepted but not applied. | + +## Migration + +### From ILP to QWP + +The existing `Sender` row API can migrate from `http::` or `tcp::` to `ws::` +or `wss::`, followed by `await sender.connect()`. QWP adds querying, replay, +store-and-forward, and richer types. Unlike ILP HTTP, a QWP `flush()` can +resolve **before** the server accepts the rows, so wait for an ACK when +committing source offsets. Legacy ILP-only keys such as `retry_timeout`, +`init_buf_size`, and `tls_ca` are rejected on QWP strings. + +### Upgrading from 4.x + +Version 5.0.0 requires Node.js 20.18.1 or newer. Existing ILP senders now +omit columns passed `null` or `undefined` (instead of throwing for most +values); `decimalColumn()` rejects non-integer scales instead of silently +coercing them. The package adds `ws` for QWP and retains the old ILP +transports. + +## ILP transports (legacy) + +Use `Sender.fromConfig("http::addr=localhost:9000;")` for ILP over HTTP, +or `tcp::addr=localhost:9009;` for TCP. ILP is ingestion-only: call +`sender.flush()` before `sender.close()` or buffered rows are lost. See the +[ILP overview](/docs/connect/compatibility/ilp/overview/) for authentication, +protocol versions, and transport configuration. ## Next steps -- [Node.js operations and reference](/docs/connect/clients/nodejs-operations/) - for pool sizing, failover, error handling, configuration, migration, and a - complete ingestion and querying example. - -- [Connect string reference](/docs/connect/clients/connect-string/) for every - configuration key. -- [Delivery semantics](/docs/concepts/delivery-semantics/) for at-least-once - delivery and deduplication. -- [Client failover](/docs/high-availability/client-failover/concepts/) and - [store-and-forward](/docs/high-availability/store-and-forward/concepts/) - concepts. -- [Query & SQL overview](/docs/query/overview/) for QuestDB SQL. -- The client's - [GitHub repository](https://github.com/questdb/nodejs-questdb-client) and the - [Community Forum](https://community.questdb.com/) for questions and issues. +- [Delivery semantics](/docs/concepts/delivery-semantics/) for replay and deduplication. +- [Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/). +- [Connect string reference](/docs/connect/clients/connect-string/) and the + [Node.js API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index d924d3260d..c7e8c807ea 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -141,7 +141,7 @@ retryable failures and stay within the failover budget. With no explicit policy, failover is enabled for established query connections but initial connection attempts are not retried. `failover=off` disables the reconnect wrapper; an explicit `egressSession.reconnect` value overrides the connect-string policy. -See [Node.js connection events](/docs/connect/clients/nodejs-operations/#connection-events). +See [Node.js connection events](/docs/connect/clients/nodejs/#connection-events). ::: @@ -379,7 +379,7 @@ start (`initial_connect_retry=async` or `lazy_connect=on`) keep retrying from startup and emit `durable-ack-unavailable`, even with `sf_dir`. With a foreground start and `sf_dir`, the first connection fails but later mismatches are retried after a successful connection. Monitor these -[connection events](/docs/connect/clients/nodejs-operations/#connection-events) and buffer +[connection events](/docs/connect/clients/nodejs/#connection-events) and buffer usage; see [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). ### Reconnect and outage handling @@ -411,8 +411,8 @@ This is the behaviour of the Java reference client and the .NET client. Other clients are aligned to it, except a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, which gives up after `reconnect_max_duration_millis`; see the -[Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect) and its -[other differences](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). +[Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect) and its +[other differences](/docs/connect/clients/nodejs/#differences-from-other-clients). If you are implementing a new client, the contract is: retry transport failures forever, surface only genuine terminal conditions, diff --git a/documentation/connect/wire-protocols/qwp-ingress-websocket.md b/documentation/connect/wire-protocols/qwp-ingress-websocket.md index 18321f7bc5..297a2cb2e5 100644 --- a/documentation/connect/wire-protocols/qwp-ingress-websocket.md +++ b/documentation/connect/wire-protocols/qwp-ingress-websocket.md @@ -1151,7 +1151,7 @@ section of the connect string reference: | Key | Default | Description | |----------------------------------|-----------|-------------------------------------------| -| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on` ([details](/docs/connect/clients/nodejs-operations/#differences-from-other-clients)). | +| `reconnect_max_duration_millis` | `300000` | Budget for the blocking sync initial connect only; the running loop retries indefinitely, except on a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on` ([details](/docs/connect/clients/nodejs/#differences-from-other-clients)). | | `reconnect_initial_backoff_millis` | `100` | First post-failure sleep. | | `reconnect_max_backoff_millis` | `5000` | Cap on per-attempt sleep. | | `initial_connect_retry` | `off` | Retry on first connect (`on`, `sync`, `async`). | diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index db808f3347..0a2bb2db11 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -83,7 +83,7 @@ both storage modes, so the `zone=` key is silently accepted on ingress connections and only takes effect on egress. The Node.js client is the exception: it applies `zone=` and `target=` to ingress too. Its other deviations are listed under -[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ### Selection priority @@ -155,7 +155,7 @@ length, and what bounds your tolerance is buffer capacity `initial_connect_retry=async`, or `lazy_connect=on`, which gives up after `reconnect_max_duration_millis` and fails with `QwpReconnectExhaustedError`; see the - [Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). + [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps reconnect storms when many producers share a cluster - Inter-host pause within a round: **none** — the client walks the full @@ -264,7 +264,7 @@ its buffer capacity, so a credential change on the cluster does not stop the producer. Monitor such a sender: the Java client reports each rejection to the sender's error handler as a retriable `SECURITY_ERROR`, and the Node.js client emits an `attempt-failed` connection event for each failed attempt. See -[Node.js connection errors](/docs/connect/clients/nodejs-operations/#connection-level-errors) +[Node.js connection errors](/docs/connect/clients/nodejs/#connection-level-errors) for the Node.js rules. Per-host credentials are outside the failover model. Use a separate connect diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index dda81c40cd..0d751ea8c0 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -19,7 +19,7 @@ first. ingress parsers accept and ignore them, so one connect string can serve both directions. The Node.js client is the exception: it applies both keys to ingress too; its other deviations are listed under -[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). They are documented in full on the [connect-string reference](/docs/connect/clients/connect-string#failover-keys); the table below summarises the failover-relevant subset. @@ -27,8 +27,8 @@ the table below summarises the failover-relevant subset. | Key | Type | Default | Notes | |---|---|---|---| | `addr` | `host:port[,host:port…]` | required | Comma-separated peer list. The two syntactic forms (`addr=h1,h2` and repeated `addr=h1;addr=h2`) accumulate. Empty entries are rejected. | -| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs-operations/#multiple-endpoints), which also ranks ingress endpoints by zone. | -| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only** (the Node.js client also applies it to ingress). Which server role the query client accepts. Other clients accept and ignore it on an ingress connect string. On the [Node.js client](/docs/connect/clients/nodejs-operations/#multiple-endpoints), set the query-side role through the typed `egress` option instead. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | +| `zone` | string | unset | Client's zone identifier (opaque, case-insensitive — `eu-west-1a`, `dc-amsterdam`, etc.). Egress prefers same-zone peers when `target` is `any` or `replica`. Silently accepted but ignored on ingress, except by the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), which also ranks ingress endpoints by zone. | +| `target` | `any` \| `primary` \| `replica` | `any` | **Egress only** (the Node.js client also applies it to ingress). Which server role the query client accepts. Other clients accept and ignore it on an ingress connect string. On the [Node.js client](/docs/connect/clients/nodejs/#multiple-endpoints), set the query-side role through the typed `egress` option instead. See [Role filter](/docs/high-availability/client-failover/concepts/#role-filter-target) for the role table. | | `auth_timeout_ms` | int (ms) | `15000` | Upper bound on the HTTP-upgrade response read per host. Does **not** cover TCP connect or TLS handshake. `connect_timeout` bounds the TCP connect separately: most clients leave it unset by default and then use the OS timeout, while Node.js defaults it to 15 s and also bounds DNS and TLS with it. On Node.js, `auth_timeout_ms` defaults to `connect_timeout` when only that key is set. Lower `auth_timeout_ms` for faster upgrade failure detection; tune the connect timeout separately. | `addr` syntax — both of these are equivalent and produce the same three-peer @@ -50,7 +50,7 @@ for the full list. The failover-relevant keys are: | Key | Type | Default | Notes | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index e162e74035..de7e60633d 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -46,7 +46,7 @@ memory mode in two. A sender in default memory mode gives up after background memory mode (`initial_connect_retry=async` or `lazy_connect=on`) retries indefinitely, as SF mode does. See the [Node.js ingestion modes](/docs/connect/clients/nodejs/#ingestion-modes) and -[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ## What "frame" means here diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 345e707e52..4a1c3158cf 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -48,7 +48,7 @@ and host-walk semantics are documented in | Key | Type | Default | Description | |---|---|---|---| -| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs-operations/#ingestion-reconnect). | +| `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | | `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 3bd147a055..6f4059043d 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -60,7 +60,7 @@ switch between them without changing application code — only the connect string. On the Node.js client, a sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, also gives up after `reconnect_max_duration_millis` of outage; see -[Differences from other clients](/docs/connect/clients/nodejs-operations/#differences-from-other-clients). +[Differences from other clients](/docs/connect/clients/nodejs/#differences-from-other-clients). ## Comparison at a glance @@ -119,7 +119,7 @@ GCS, or NFS). even with `sf_dir`. With `sf_dir` and a foreground start, the first connection fails but later mismatches are retried after a successful connection. Retrying Node.js senders emit `durable-ack-unavailable` - [connection events](/docs/connect/clients/nodejs-operations/#connection-events). + [connection events](/docs/connect/clients/nodejs/#connection-events). Monitor retrying senders and their buffer usage: continued buffering can fill the journal or memory queue even though startup succeeded. See [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). diff --git a/documentation/query/overview.md b/documentation/query/overview.md index 71317d0433..ee3723ef11 100644 --- a/documentation/query/overview.md +++ b/documentation/query/overview.md @@ -121,7 +121,7 @@ SQL writes such as `INSERT`. Each client describes its failover behavior: [Rust](/docs/connect/clients/rust/#failover-and-errors), [C and C++](/docs/connect/clients/c-and-cpp/#querying-data), [.NET](/docs/connect/clients/dotnet/#failover), and -[Node.js](/docs/connect/clients/nodejs-operations/#query-failover). +[Node.js](/docs/connect/clients/nodejs/#query-failover). The Rust, C++, and Python clients hand back results as Arrow record batches. That is the native memory layout of diff --git a/documentation/sidebars.js b/documentation/sidebars.js index 1b91bfb023..e824810980 100644 --- a/documentation/sidebars.js +++ b/documentation/sidebars.js @@ -78,11 +78,6 @@ module.exports = { type: "doc", label: "Node.js", }, - { - id: "connect/clients/nodejs-operations", - type: "doc", - label: "Node.js operations and reference", - }, { id: "connect/clients/c-and-cpp", type: "doc", diff --git a/plugins/raw-markdown/convert-components.test.js b/plugins/raw-markdown/convert-components.test.js index adf9ab35ca..70ccd63aa3 100644 --- a/plugins/raw-markdown/convert-components.test.js +++ b/plugins/raw-markdown/convert-components.test.js @@ -60,5 +60,5 @@ test('preserves imports from the Node.js Quick start while removing its MDX impo const result = removeImports(content) assert.doesNotMatch(result, /^import SfDedupWarning from /m) - assert.match(result, /```typescript\nimport \{\n connectQwpNodeClient,\n QwpEgressQueryError,\n\} from "@questdb\/nodejs-client";/) + assert.match(result, /```typescript\nimport \{ connectQwpNodeClient \} from "@questdb\/nodejs-client";/) }) diff --git a/shared/clients.json b/shared/clients.json index 86d54ac4d8..393030b9df 100644 --- a/shared/clients.json +++ b/shared/clients.json @@ -112,7 +112,7 @@ "protocol": "PGWire" }, { - "href": "/docs/connect/clients/nodejs-operations#ilp-transports-legacy", + "href": "/docs/connect/clients/nodejs#ilp-transports-legacy", "name": "Node.js", "description": "Node.js client for ILP ingestion over HTTP and TCP, alongside QWP.", From 25903c9ccfc1270012b5cce93936f063cb53e70e Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 00:26:42 +0100 Subject: [PATCH 19/25] docs: clarify QWP delivery and Node.js recovery guidance --- documentation/concepts/delivery-semantics.md | 31 ++++++----- .../connect/clients/connect-string.md | 14 ++++- documentation/connect/clients/nodejs.md | 55 ++++++++++++++++++- .../store-and-forward/configuration.md | 2 +- 4 files changed, 82 insertions(+), 20 deletions(-) diff --git a/documentation/concepts/delivery-semantics.md b/documentation/concepts/delivery-semantics.md index 201651a5a4..a72b78230a 100644 --- a/documentation/concepts/delivery-semantics.md +++ b/documentation/concepts/delivery-semantics.md @@ -2,24 +2,25 @@ title: Delivery semantics sidebar_label: Delivery semantics description: - How QuestDB clients deliver data (at-least-once), where duplicate rows can - arise, and how to combine designated timestamps with deduplication for - exactly-once outcomes. + How QuestDB QWP/WebSocket senders replay unacknowledged writes, where + duplicates arise, and how to use deduplication for exactly-once outcomes. --- -QuestDB clients deliver data **at-least-once**: a sender keeps every row your -application publishes until the server acknowledges it, and resends it after a -failure, so under failure a row may arrive more than once. Storing each row +QuestDB QWP/WebSocket senders retry **published, unacknowledged batches** +while they retain them. This gives at-least-once delivery under transport +failures, but a row may be stored more than once. QWP/UDP is fire-and-forget: +it has no acknowledgement or retry, so rows can be lost. Storing each row exactly once is the application's responsibility, and QuestDB provides the mechanisms to make it routine. -The guarantee holds while the sender runs. Without -[store-and-forward](/docs/high-availability/store-and-forward/concepts/), +Without [store-and-forward](/docs/high-availability/store-and-forward/concepts/), unacknowledged rows live in memory and are lost if the process exits, or the sender closes, before the server acknowledges them. A Node.js sender in default memory mode also gives up after `reconnect_max_duration_millis` of outage; see the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). +Server rejections and exhausted buffer capacity can also stop delivery; handle +those errors rather than assuming every attempted write will arrive. This page explains where duplicates come from and how to suppress them. @@ -27,14 +28,14 @@ This page explains where duplicates come from and how to suppress them. | Property | Meaning | Where it comes from | |----------|---------|---------------------| -| **At-most-once** | Each row reaches the server zero or one times. Rows can be lost. | A "fire and forget" client that does not retransmit on failure. | -| **At-least-once** | Each row reaches the server one or more times. No row is lost; duplicates are possible. | A client that retransmits unacknowledged data after a transport error. **QuestDB clients provide this while the sender runs**, and across restarts with store-and-forward. | +| **At-most-once** | Each row reaches the server zero or one times. Rows can be lost. | Fire-and-forget delivery without ACKs or retries, such as QWP/UDP. | +| **At-least-once** | Each row reaches the server one or more times. Duplicates are possible. | QWP/WebSocket senders retry published batches while they retain them; disk-backed store-and-forward extends replay across process restarts. | | **Exactly-once** | Each row is stored exactly once. | At-least-once delivery plus server-side deduplication on a key covering row identity. | -QuestDB's clients retransmit unacknowledged batches after transport errors, -host failovers, and process restarts. The trade-off is deliberate: losing -data silently is the worse failure mode. The cost is that the application -must tolerate or suppress duplicates. +QWP/WebSocket senders retransmit unacknowledged batches after transport errors +and host failovers, and across process restarts with store-and-forward. The +trade-off is deliberate: losing data silently is the worse failure mode. The +cost is that the application must tolerate or suppress duplicates. ## Where duplicates come from @@ -47,7 +48,7 @@ the server confirms a batch, the client reconnects and re-sends. If the server had already committed the batch but the acknowledgement was lost in flight, the second send produces duplicates. -This path applies to every QuestDB client deployment. +This path applies to QWP/WebSocket senders, not QWP/UDP. ### Multi-host failover replay diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 9b83453ad8..f7c9e3fc7e 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -132,10 +132,22 @@ wss::addr=questdb.example.com:443;username=admin;password=secret; ### Production with a custom trust store +For Node.js and clients that accept PEM roots, use a PEM file without a password: + +```text +wss::addr=questdb.example.com:443;username=admin;password=secret;tls_roots=/etc/questdb/ca.pem; ``` -wss::addr=questdb.example.com:443;username=admin;password=secret;tls_roots=/etc/questdb/ca-roots;tls_roots_password=changeit; + +For clients that accept password-protected JKS or PKCS#12 stores, supply the +password as well: + +```text +wss::addr=questdb.example.com:443;username=admin;password=secret;tls_roots=/etc/questdb/ca-roots.p12;tls_roots_password=changeit; ``` +Node.js rejects `tls_roots_password` and accepts only PEM roots; Go accepts +neither key and uses the OS trust store. See [TLS](#tls) for formats by client. + ### Ingest with store-and-forward across multiple nodes ``` diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 164769205d..92eea32988 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -407,6 +407,10 @@ try { } ``` +If this client closes before an ACK, the journal keeps the unacknowledged row. +After QuestDB is reachable again, reopen the same `sf_dir` and `sender_id` to +replay it before checking query visibility; see [Read-after-write](#read-after-write). + `sf_durability=memory` (the default) survives a process crash, not a power failure; `periodic` checkpoints and `append` syncs each append. The sender's first connection is foreground by default; adding `sf_dir` alone does not @@ -540,14 +544,27 @@ outcome before retrying, or make the statement idempotent. An ACK confirms commitment to the WAL, **not** query visibility: WAL apply is asynchronous. Create the table before writing, then poll for a stable event -ID with a deadline. For the `trades_sf` table above, after publishing -`trade-12345`: +ID with a deadline. If the store-and-forward example above closed before its +batch was acknowledged, run this example **after QuestDB is reachable again**. +It reopens the same journal slot, waits for the pending frames to be +acknowledged, and only then polls for `trade-12345`. A query-only client without +that journal would not replay the row. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;sf_dir=./questdb-sf;sender_id=trades;" + + "sf_max_segment_bytes=1m;failover=off;", +); try { + const sender = await db.borrowSender(); + try { + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); + } finally { + await sender.close(); + } + const lease = await db.borrowQuery(); try { const deadline = Date.now() + 10_000; @@ -681,6 +698,38 @@ connect string also filters ingestion**, so use the typed query-only option "prefer replica and fall back to primary". If no replica is up at startup, set `query_pool_min=0` to defer the query connection. +For a query-only client with only replica endpoints, set `sender_pool_min=0`: +otherwise the default pool prewarms a sender and startup fails because no +primary can accept its write connection. Set the replica filter in the typed +query options so it applies only to queries. If you later borrow a sender, +include a primary endpoint in `addr`. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient( + "wss::addr=db-a.example.com:9000,db-b.example.com:9000;" + + "token=YOUR_TOKEN;sender_pool_min=0;", + { egress: { target: "replica" } }, +); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol FROM trades LIMIT 10", + ); + for await (const batch of query) { + for (const row of batch.rows()) console.log(row); + } + await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + ### Ingestion reconnect Senders resend unacknowledged batches after a disconnect. A sender in default diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index 4a1c3158cf..cacfd3bf6f 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -65,7 +65,7 @@ Opt in to object-store-durable trim. See | Key | Type | Default | Description | |---|---|---|---| | `request_durable_ack` | bool | `off` | Opt-in via the upgrade header `X-QWP-Request-Durable-Ack: true`. Trim is then driven by `STATUS_DURABLE_ACK` frames only; OK frames no longer advance the trim watermark. A missing `X-QWP-Durable-Ack: enabled` echo is terminal in most clients; Java and some Node.js senders keep retrying (see [Concepts](/docs/high-availability/store-and-forward/concepts/#trim-how-unacked-data-is-reclaimed)). WebSocket transports only. | -| `durable_ack_keepalive_interval_millis` | int (ms) | `200` | Cadence of WebSocket PING the I/O loop sends while there are pending durable confirmations and the producer is idle. `0` or negative disables. | +| `durable_ack_keepalive_interval_millis` | int (ms) | `200` | Cadence of WebSocket PING while durable confirmations are pending and the producer is idle. `0` disables the PING; some clients also accept negative values. Node.js rejects negative values, and explicitly setting even `0` requests durable ACK, so an unsupported server may reject the connection. See [Node.js durable acknowledgement](/docs/connect/clients/nodejs/#durable-acknowledgement). | ## Error-handling keys From e9a12cbaf1cd70f4de70e6d1e7f4d3b8a88c051f Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 12:37:58 +0100 Subject: [PATCH 20/25] docs: address review findings in the Node.js client guide - Make Read-after-write self-contained and move journal replay to its own subsection, which polls for the row instead of waiting on recovered ACKs - Correct the batch-size advice and the typed store-and-forward defaults - Document the result row shape, failover defaults, the full bind setter table, cancellation cost without a credit window, and shutdown cancels - Expand error handling: typed catch examples, sender error fields and policies, upgrade error kinds, 401/403 retry rules, and a runbook for a journal blocked by a terminal rejection - List all connection events and warn that a typed egress reconnect object re-enables query failover - Add examples for the standalone Sender, compiled writers, offset commits, transactions, durable ACK, UDP, DDL/DML, cancellation, result views, and query failover; restore a runnable ILP example - Consolidate startup and outage modes into one section and complete the Differences table - Note Node.js full jitter on the failover and store-and-forward pages and stop listing every server rejection as terminal --- documentation/connect/clients/nodejs.md | 902 ++++++++++++++---- .../client-failover/concepts.md | 5 +- .../client-failover/configuration.md | 2 +- .../store-and-forward/configuration.md | 2 +- 4 files changed, 740 insertions(+), 171 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 92eea32988..2949a3d3c2 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -104,43 +104,39 @@ unknown keys and duplicate keys fail validation. Escape a semicolon in a value as `;;`. See the [connect string reference](/docs/connect/clients/connect-string/) for the shared keys and [Differences from other clients](#differences-from-other-clients) -for Node.js exceptions. +for Node.js exceptions. A second argument takes typed options; see +[Programmatic options](#programmatic-options). ### Standalone Sender -For ingestion without a pool, use -`const sender = await Sender.fromConfig("ws::addr=localhost:9000;")`, then -`await sender.connect()` before writing and `await sender.close()` in a -`finally` block. The standalone `Sender` also supports ILP (`http::` and -`tcp::`), but exposes fewer fluent QWP column methods; use `sender.writer()` -for other types, or a pooled sender. For a typed standalone QWP sender, use -`await connectQwpNodeSender({ url: "ws://localhost:9000/write/v4" })`. - -### Programmatic options - -The second argument of `connectQwpNodeClient` accepts typed `sender`, -`ingressSession`, `egressSession`, `egress`, `webSocket`, `storeAndForward`, and -`pool` options. For example: +For ingestion without a pool, create a `Sender` from a connect string, call +`connect()` before writing, and close it in a `finally` block: ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { Sender } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { - sender: { awaitServerAck: true }, - ingressSession: { - onSenderError: (error) => - console.error("rejected batch:", error.category, error.serverMessage), - }, - pool: { senderPoolMax: 2 }, -}); -await db.close(); +const sender = await Sender.fromConfig("ws::addr=localhost:9000;"); +try { + await sender.connect(); + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .floatColumn("price", 2615.54) // DOUBLE: the Sender has no doubleColumn() + .floatColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); +} finally { + await sender.close(); +} ``` -Typed options override the corresponding connect-string settings. A typed -`ingressSession.reconnect` or `egressSession.reconnect` object **replaces** -the entire policy from the string, not just the fields specified. For the -full typed API, see the -[client reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). +The standalone `Sender` has only the nine column methods it shares with ILP +(see [Column methods](#column-methods)); use `sender.writer()` for other types, +or a pooled sender. It also speaks ILP over `http::` and `tcp::`; see +[ILP transports](#ilp-transports-legacy). For a typed standalone QWP sender +with every column method, use +`await connectQwpNodeSender({ url: "ws://localhost:9000/write/v4" })`. @@ -164,8 +160,13 @@ Basic auth uses `username=...;password=...;`. With `wss`, the client checks certificates against Node.js's bundled CA store. For a private CA, set `tls_roots=/path/to/ca.pem` or `NODE_EXTRA_CA_CERTS`; only PEM roots are supported. `tls_verify=unsafe_off` disables verification for development -and cannot be combined with `tls_roots`. A token is read when the client is -created: create a new client to rotate it. +and cannot be combined with `tls_roots`. + +| Path | Status | Alternative | +|---|---|---| +| OIDC token acquisition or refresh | Not supported. The client does not talk to an identity provider. | Get an access token from your identity provider, pass it as `token`, and create a new client before it expires. See [OpenID Connect](/docs/security/oidc/). | +| Client certificates (mTLS) | Not supported. QuestDB does not negotiate client certificates. | Use a token or basic auth over `wss`. | +| Token rotation on a running client | Not supported. Every connection, including reconnects, sends the token the client was created with. | Close the client and create a new one with the new token. For a token rejected on reconnect, see [Connection-level errors](#connection-level-errors). | ## The connection pool @@ -177,15 +178,32 @@ pool**, but does not normally wait for their acknowledgements. A returned lease or sender must not be reused. Size pools to the number of simultaneous borrows, and close the client on shutdown. -### Starting while QuestDB is down - -`connectQwpNodeClient()` normally connects at startup. `lazy_connect=on` -starts senders in the background, forces `query_pool_min=0`, and lets the -client start while QuestDB is down. A query borrowed before QuestDB is -reachable can still fail. Memory-only rows are lost if the process exits; -add `sf_dir` to keep them on disk. A locked journal fails startup even with -`lazy_connect=on`. With `sf_dir` and a background start, retries begin from -startup; without a background start, the first connection must succeed. +### Startup and outage modes {#ingestion-modes} + +Two independent choices decide what a sender does while QuestDB is +unreachable: + +- **Startup.** A *foreground start*, the default, connects inside + `connectQwpNodeClient()`, which rejects if QuestDB is down. A *background + start* (`lazy_connect=on`) returns immediately and connects senders in the + background. `lazy_connect=on` sets `query_pool_min` to 0 and rejects a + positive value, and a query borrowed before QuestDB is reachable fails. A + standalone `Sender` gets a background start with + `initial_connect_retry=async`. +- **Storage.** Without `sf_dir`, unacknowledged rows live in memory and are + lost if the process exits. With `sf_dir`, they are journaled to disk and + replayed after a restart; see [Store-and-forward](#store-and-forward). + +| Mode | Enabled by | QuestDB down at startup | During an outage | +|---|---|---|---| +| Default memory | Neither `sf_dir` nor a background start | Startup rejects | `flush()`, and `at()` when it triggers an auto-flush, wait for the reconnect for up to `reconnect_max_duration_millis` (5 minutes by default). Then the sender fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost | +| Background memory | A background start without `sf_dir` | Starts; rows queue in memory | Rows queue in memory, up to `sf_max_total_bytes` (128 MiB by default); retries continue indefinitely | +| Store-and-forward | `sf_dir`, with either startup | Foreground: startup rejects. Background: starts | Rows go to the disk journal; retries continue indefinitely, from startup with a background start or after the first successful connection otherwise | + +A foreground start fails fast unless `initial_connect_retry=on` or any +`reconnect_*` key is set; then it retries the first connection for up to +`reconnect_max_duration_millis` before rejecting. A locked journal fails +startup even with a background start. ### Closing the pooled client @@ -289,21 +307,48 @@ arrays as `{ dimensions, values }`. For a stream of objects with a fixed shape, compile a writer once with `sender.writer(table, schema)` and call `writer.row(object)` or -`writer.rows(iterable)`. Schema builders include `symbol()`, `double()`, -`varchar()`, and `designatedTimestamp("ms")`. The writer validates each row; -its `QwpWriterRowError` names the offending table, column, and row index. -See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) -for all builders. +`writer.rows(iterable)`. The package exports the schema builders, such as +`symbol()`, `double()`, `varchar()`, and `designatedTimestamp("ms")`. Each +schema key names a column, except the `designatedTimestamp()` key: it supplies +the row's designated timestamp (named `timestamp` on a new table) and is +required in every row. -### Ingestion modes {#ingestion-modes} +```typescript +import { + connectQwpNodeClient, + designatedTimestamp, + double, + symbol, +} from "@questdb/nodejs-client"; -Storage (`sf_dir`) and the initial-connect choice are independent: +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + const trades = sender.writer("trades", { + symbol: symbol(), + side: symbol(), + price: double(), + amount: double(), + timestamp: designatedTimestamp("ms"), + }); + await trades.rows([ + { symbol: "ETH-USD", side: "buy", price: 2615.54, amount: 0.5, timestamp: Date.now() }, + { symbol: "BTC-USD", side: "sell", price: 61234.5, amount: 0.01, timestamp: Date.now() }, + ]); + await sender.flush(); + } finally { + await sender.close(); // the writer cannot be used after its sender is closed + } +} finally { + await db.close(); +} +``` -| Mode | Enabled by | During an outage | -|---|---|---| -| Default memory | Neither `sf_dir` nor a background start | `flush()` waits for reconnect, for up to 5 minutes by default | -| Background memory | No `sf_dir`; `lazy_connect=on` or `initial_connect_retry=async` | Rows queue in memory; retries continue | -| Store-and-forward | `sf_dir`, with either startup choice | Rows go to the disk journal; retries continue after the first successful connection, or from startup with a background start | +The writer validates each row; its `QwpWriterRowError` names the offending +table, column, and row index. See the +[client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +for all builders. ### Flushing @@ -326,10 +371,13 @@ retry `flush()`; do not write the rows again. Configure the cap with #### Batch size limits QuestDB advertises its maximum batch size on connection (about 2 MiB on a -default server). A batch too large to send fails with `QwpBatchTooLargeError`; -call `reset()` and rebuild it in smaller batches. If a sender starts offline, -it cannot yet know the server limit. With `sf_dir` and `lazy_connect=on`, set -`sf_max_segment_bytes=1m` so an oversized batch cannot block journal replay. +default server). The client splits a larger batch into several frames at row +boundaries. A single row larger than the limit fails with +`QwpBatchTooLargeError`: call `reset()` and shrink that row, for example a +large VARCHAR or BINARY value. A sender that starts offline cannot yet know +the server limit. With `sf_dir` and `lazy_connect=on`, set +`sf_max_segment_bytes=1m`: it also caps each journaled frame at 1 MiB, so an +oversized frame cannot block journal replay. ### Awaiting acknowledgements @@ -348,9 +396,56 @@ already published them. To make each flush wait, use typed When consuming from Kafka or another source, record each batch's `publishedSequence` with its last source offset. Commit only the newest offset -whose sequence is at or below `acknowledgedSequence`. Request -[durable acknowledgement](#durable-acknowledgement) if the offset must also -survive a primary failure. +whose sequence is at or below `acknowledgedSequence`: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +interface Fill { offset: bigint; symbol: string; price: number; amount: number; tsMs: number } +// Stand-ins for your consumer: replace them with your Kafka client's calls. +const batches: Fill[][] = [ + [{ offset: 41n, symbol: "ETH-USD", price: 2615.54, amount: 0.5, tsMs: Date.now() }], + [{ offset: 42n, symbol: "BTC-USD", price: 61234.5, amount: 0.01, tsMs: Date.now() }], +]; +const commitOffset = async (offset: bigint) => console.log("commit", offset); + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + const pending: { sequence: bigint; offset: bigint }[] = []; + const commitAcknowledged = async () => { + let newest: bigint | undefined; + while (pending.length > 0 && pending[0].sequence <= sender.acknowledgedSequence) { + newest = pending.shift()!.offset; + } + if (newest !== undefined) await commitOffset(newest); + }; + try { + for (const batch of batches) { + for (const fill of batch) { + await sender + .table("trades") + .symbol("symbol", fill.symbol) + .doubleColumn("price", fill.price) + .doubleColumn("amount", fill.amount) + .at(fill.tsMs, "ms"); + } + await sender.flush(); + pending.push({ sequence: sender.publishedSequence, offset: batch[batch.length - 1].offset }); + await commitAcknowledged(); + } + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); + await commitAcknowledged(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +Request [durable acknowledgement](#durable-acknowledgement) if the offset must +also survive a primary failure. ### Transactions @@ -361,13 +456,39 @@ table, not across tables, and QuestDB can commit early when the table exceeds Closing a standalone sender without `flush()` rolls back the open transaction; returning a pooled sender with `close()` flushes and commits instead. +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;transaction=on;"); +try { + const sender = await db.borrowSender(); + try { + for (const [side, price] of [["buy", 2615.54], ["sell", 2615.62]] as const) { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", side) + .doubleColumn("price", price) + .doubleColumn("amount", 0.5) + .at(Date.now(), "ms"); + } + await sender.commit(); // both rows become visible together + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + ### Store-and-forward -Set `sf_dir` to journal batches across process restarts. For an offline start, -also set `lazy_connect=on` (or `initial_connect_retry=async` on a standalone -sender). Create a deduplicated table **before** ingestion if duplicates are -unacceptable: a missing table is auto-created without DEDUP. Keep event IDs -and timestamps stable across retries. +Set `sf_dir` to journal batches across process restarts. To start while +QuestDB is down, add a background start (`lazy_connect=on`), as below; see +[Startup and outage modes](#ingestion-modes). Create a deduplicated table +**before** ingestion if duplicates are unacceptable: a missing table is +auto-created without DEDUP. Keep event IDs and timestamps stable across +retries. @@ -407,19 +528,64 @@ try { } ``` -If this client closes before an ACK, the journal keeps the unacknowledged row. -After QuestDB is reachable again, reopen the same `sf_dir` and `sender_id` to -replay it before checking query visibility; see [Read-after-write](#read-after-write). +If this client closes before an ACK, the journal keeps the unacknowledged row +until a client reopens the same `sf_dir` and `sender_id`; see +[Replaying the journal after a restart](#replaying-the-journal-after-a-restart). `sf_durability=memory` (the default) survives a process crash, not a power -failure; `periodic` checkpoints and `append` syncs each append. The sender's -first connection is foreground by default; adding `sf_dir` alone does not -make it lazy. With `sf_dir` plus `lazy_connect=on`, it retries from startup. +failure; `periodic` checkpoints and `append` syncs each append. `sf_dir` alone +keeps a foreground start; see [Startup and outage modes](#ingestion-modes). A terminally rejected batch stays at the head of the journal and blocks later -rows until fixed. See the +rows until fixed; see +[Recovering from a terminal rejection](#recovering-from-a-terminal-rejection). +See also the [store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/) and [operating guide](/docs/high-availability/store-and-forward/operating-and-tuning/). +#### Replaying the journal after a restart + +Run a client with the same `sf_dir` and `sender_id` once QuestDB is reachable +again. The pool's first sender reopens the journal slot (`trades-0` here) at +startup and replays the unacknowledged frames in the background, so keep the +default `sender_pool_min=1`. Then poll for a row you know was written: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;sf_dir=./questdb-sf;sender_id=trades;" + + "sf_max_segment_bytes=1m;failover=off;", +); +try { + const lease = await db.borrowQuery(); + try { + const deadline = Date.now() + 10_000; + let visible = false; + while (!visible && Date.now() < deadline) { + const query = await lease.query( + "SELECT trade_id FROM trades_sf WHERE trade_id = $1 LIMIT 1", + { + binds: (binds) => binds.setVarchar(0, "trade-12345"), + timeoutMs: Math.max(1, deadline - Date.now()), + }, + ); + for await (const batch of query) visible ||= batch.rowCount > 0; + await query.completion; + if (!visible) await new Promise((resolve) => setTimeout(resolve, 100)); + } + if (!visible) throw new Error("trade not visible in time"); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`failover=off` prevents a replayed query from invalidating the result +mid-poll. A fixed sleep without checking for the row is not a visibility +guarantee. + #### Journal capacity {#sf-capacity} `sf_max_total_bytes=10g` is a target, not a hard disk quota: transaction @@ -442,12 +608,39 @@ Do not let Node.js and another client's OS-lock-based sender use the same On QuestDB Enterprise with replication, `request_durable_ack=on` makes the acknowledgement watermark wait until the WAL has been uploaded to object -storage. `sender: { awaitDurableAck: true }` also makes each `flush()` wait. -If the server lacks support, a foreground first connection fails with +storage. `sender: { awaitDurableAck: true }` also makes each `flush()` wait: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); +const db = await connectQwpNodeClient( + `wss::addr=db.example.com:9000;token=${token};request_durable_ack=on;`, + { sender: { awaitDurableAck: true } }, +); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .doubleColumn("price", 2615.54) + .at(Date.now(), "ms"); + await sender.flush(); // resolves once the batch is in object storage + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +If the server lacks support, a foreground start fails with `QwpDurableAckUnavailableError` (wrapped in `QwpPoolResourceError` when -pooled). A background-started sender retries from startup and emits +pooled). A sender with a background start retries from startup and emits `durable-ack-unavailable` connection events **even with `sf_dir`**. With -`sf_dir` and foreground startup, only later mismatches, after a successful +`sf_dir` and a foreground start, only later mismatches, after a successful connection, are retried. Monitor these events and journal capacity: a successful background start does not prove durable ACK is available. @@ -459,6 +652,25 @@ fire-and-forget ingestion. Enable the server's TLS, auth, ACK, retries, transactions, or store-and-forward; use WebSocket for reliable writes. +```typescript +import { Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig("udp::addr=localhost:9007;"); +try { + await sender.connect(); + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.5) + .at(Date.now(), "ms"); + await sender.flush(); // sends datagrams; nothing confirms delivery +} finally { + await sender.close(); +} +``` + ## Querying Borrow one query lease per concurrent query. A lease executes one query at a @@ -482,7 +694,9 @@ try { }, ); for await (const batch of query) { - for (const row of batch.rows()) console.log(row); + for (const [timestamp, symbol, price] of batch.rows()) { + console.log(new Date(Number((timestamp as bigint) / 1000n)), symbol, price); + } } await query.completion; } finally { @@ -496,7 +710,14 @@ try { `lease.query()` returns a handle with async result batches and a `completion` promise. For DDL/DML, await `completion` without iterating. A batch has `rowCount`, `columns`, `rows()`, `get(rowIndex, columnIndex)`, and -`batchSequence`. A query can fail during iteration as well as at completion. +`batchSequence`. `rows()` yields one array per row, in SELECT column order: +destructure or index it, because a row has no named properties (`row.price` +is `undefined`). `batch.columns[i].name` gives the column names. A query can +fail during iteration as well as at completion. + +Query failover is on by default: if the connection is lost, the query runs +again from its first row, so code that accumulates rows must handle the +restart. See [Query failover](#query-failover). ### Reading result values @@ -514,52 +735,103 @@ promise. For DDL/DML, await `completion` without iterating. A batch has Other types include IPv4 (signed 32-bit `number`), GEOHASH and LONG256 objects. `JSON.stringify()` cannot serialize `bigint`: convert it to a string -first. Current servers cannot return INTERVAL, an untyped NULL, or DECIMAL +first. Convert a TIMESTAMP to a `Date` with `new Date(Number(value / 1000n))`. +Current servers cannot return INTERVAL, an untyped NULL, or DECIMAL with precision 9 or less over QWP; cast those in SQL to a supported type. ### Bind parameters -Bind indexes start at 0 for `$1` and must be set in ascending order without -gaps. Use `setVarchar`, `setInt`, `setLong`, `setDouble`, -`setTimestampMicros`, `setTimestampNanos`, `setUuid`, or other typed setters -on the `binds` callback. Use `setNull(index, QWP_COLUMN_TYPE.DOUBLE)` for a -typed NULL; BINARY, IPv4, and arrays have no direct bind setter. +Set bind values in the `binds` callback of `query()`. Index 0 binds `$1`, and +indexes must be set in ascending order without gaps. A placeholder after the +last one you set is treated as NULL rather than rejected, so set every +placeholder. + +| QuestDB type | Setter | JavaScript value | +|---|---|---| +| BOOLEAN | `setBoolean(i, value)` | `boolean` | +| BYTE, SHORT, INT | `setByte`, `setShort`, `setInt` | `number` | +| LONG | `setLong(i, value)` | `bigint` or safe-integer `number` | +| FLOAT, DOUBLE | `setFloat`, `setDouble` | `number` | +| CHAR | `setChar(i, value)` | one-character `string` | +| VARCHAR, SYMBOL | `setVarchar(i, value)` | `string` or `null`; SYMBOL has no setter of its own | +| TIMESTAMP | `setTimestampMicros(i, value)` | microseconds, such as `BigInt(date.getTime()) * 1000n` | +| TIMESTAMP_NS | `setTimestampNanos(i, value)` | nanoseconds as a `bigint` | +| DATE | `setDate(i, value)` | milliseconds, such as `date.getTime()` | +| UUID | `setUuid(i, value)` or `setUuid(i, low, high)` | canonical UUID `string`, or two 64-bit halves | +| DECIMAL | `setDecimal64(i, scale, unscaled)`, `setDecimal128(i, scale, low, high)`, `setDecimal256(i, scale, lowLow, lowHigh, highLow, highHigh)` | unscaled value in 64-bit parts, after the scale | +| LONG256 | `setLong256(i, word0, word1, word2, word3)` | four 64-bit words | +| GEOHASH | `setGeohash(i, precisionBits, value)` | geohash bits as an integer | + +For a typed NULL, use `setNull(i, QWP_COLUMN_TYPE.DOUBLE)` (import +`QWP_COLUMN_TYPE` from the package), `setNullDecimal64/128/256(i, scale)`, or +`setNullGeohash(i, precisionBits)`. The decimal setters take the scale +**before** the unscaled value, the reverse of +`decimal64Column(name, unscaled, scale)`. BINARY, IPv4, arrays, and INTERVAL +have no bind setter. ### DDL and DML statements `CREATE`, `ALTER`, `DROP`, `TRUNCATE`, `INSERT`, and `UPDATE` use `query()`. For a statement without result batches, `completion.kind` is `"exec-done"`. -Only `INSERT` reliably provides a row count in `rowsAffected`; a WAL -`UPDATE` can report a transaction number instead. +Only `INSERT` reliably provides a row count in `rowsAffected`, a `bigint`; a +WAL `UPDATE` can report a transaction number instead. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// failover=off: a lost connection fails the INSERT instead of running it again. +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "INSERT INTO trades (timestamp, symbol, side, price, amount) " + + "VALUES (now(), 'ETH-USD', 'buy', 2615.54, 0.5)", + ); + const completion = await query.completion; + if (completion.kind === "exec-done") { + console.log(`inserted ${completion.rowsAffected} row(s)`); + } + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` :::warning SQL writes can run twice -With `failover=on`, a lost connection can re-execute in-flight SQL, including -`INSERT`. Use `failover=off` for non-idempotent SQL and check an uncertain -outcome before retrying, or make the statement idempotent. +Query failover is on by default, so a lost connection can re-execute in-flight +SQL, including `INSERT`. Use `failover=off` for non-idempotent SQL and check an +uncertain outcome before retrying, or make the statement idempotent. A typed +`egressSession.reconnect` object turns failover back on even when the connect +string says `failover=off`; see [Typed reconnect policy](#typed-reconnect-policy). ::: ### Read-after-write An ACK confirms commitment to the WAL, **not** query visibility: WAL apply -is asynchronous. Create the table before writing, then poll for a stable event -ID with a deadline. If the store-and-forward example above closed before its -batch was acknowledged, run this example **after QuestDB is reachable again**. -It reopens the same journal slot, waits for the pending frames to be -acknowledged, and only then polls for `trade-12345`. A query-only client without -that journal would not replay the row. +is asynchronous. Create the table before writing, wait for the ACK, then poll +for the row with a deadline: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient( - "ws::addr=localhost:9000;sf_dir=./questdb-sf;sender_id=trades;" + - "sf_max_segment_bytes=1m;failover=off;", -); +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); try { + const eventTime = Date.now(); // a stable key to find the row again const sender = await db.borrowSender(); try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .doubleColumn("price", 2615.62) + .doubleColumn("amount", 0.25) + .at(eventTime, "ms"); + await sender.flush(); await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } finally { await sender.close(); @@ -571,9 +843,10 @@ try { let visible = false; while (!visible && Date.now() < deadline) { const query = await lease.query( - "SELECT trade_id FROM trades_sf WHERE trade_id = $1 LIMIT 1", + "SELECT price FROM trades WHERE timestamp = $1 AND symbol = $2 LIMIT 1", { - binds: (binds) => binds.setVarchar(0, "trade-12345"), + binds: (binds) => + binds.setTimestampMicros(0, BigInt(eventTime) * 1000n).setVarchar(1, "ETH-USD"), timeoutMs: Math.max(1, deadline - Date.now()), }, ); @@ -581,7 +854,7 @@ try { await query.completion; if (!visible) await new Promise((resolve) => setTimeout(resolve, 100)); } - if (!visible) throw new Error("trade not visible in time"); + if (!visible) throw new Error("row not visible in time"); } finally { await lease.close(); } @@ -592,30 +865,103 @@ try { `failover=off` prevents a replayed query from invalidating the result mid-poll. A fixed sleep without checking for the row is not a visibility guarantee. +After a store-and-forward restart, see +[Replaying the journal after a restart](#replaying-the-journal-after-a-restart). ### Cancellation and timeouts Set `timeoutMs` per query (or `egressSession.queryTimeoutMs` by default). A deadline cancels the query and reports `QwpEgressQueryTimeoutError`. -Leaving a `for await` loop early also starts cancellation. `query.cancel()` -requests cancellation but does not wait for it; returning the lease waits for -the cancellation to drain up to `query_close_timeout_ms` (5 seconds by -default), then discards the connection if needed. +Leaving a `for await` loop early also cancels the query, and `completion` +then rejects with `QwpEgressQueryAbandonedError`. `query.cancel()` requests +cancellation but does not wait for it. + +Cancellation is prompt only with a credit window (see +[Flow control](#flow-control)). Without one, the server keeps streaming after +a cancel: returning the lease waits up to `query_close_timeout_ms` (5 seconds +by default) for the stream to drain, can take about twice that in total, and +then discards the connection. Set `initialCredit` on any query you may time +out, cancel, or stop early: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT timestamp, symbol, price FROM trades", { + timeoutMs: 5_000, + initialCredit: 1024 * 1024, + }); + let seen = 0; + let stoppedEarly = false; + for await (const batch of query) { + seen += batch.rowCount; + if (seen >= 10_000) { + stoppedEarly = true; + break; // cancels the rest of the result + } + } + // After a break, completion rejects with QwpEgressQueryAbandonedError. + if (!stoppedEarly) await query.completion; + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` ### Flow control -Without a credit window, a slow consumer can buffer a large result in memory. -Set `initialCredit: 1024 * 1024` on a query, as above, or set -`initial_credit` in the connect string. The client replenishes credit as your -loop consumes batches. Use `autoCredit: false` and `query.grantCredit(bytes)` -for manual control. +Without a credit window, the server streams as fast as it can: a slow consumer +can buffer a large result in memory, and cancelling the query is slow. Set +`initialCredit: 1024 * 1024` on a query, as above, or `initial_credit=1048576` +in the connect string; `initial_credit` takes plain bytes, not a size suffix +such as `1m`. The client replenishes credit as your loop consumes batches. Use +`autoCredit: false` and `query.grantCredit(bytes)` for manual control. ### Zero-copy result views For hot paths, `lease.queryViews(sql, callback)` reads typed values directly from received bytes instead of materializing arrays. Views and byte slices are valid only until the callback returns; copy them if you need to retain -them. See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). +them. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + let notional = 0; + const query = await lease.queryViews( + "SELECT price, amount FROM trades WHERE symbol = 'ETH-USD'", + (batch) => { + const price = batch.column(0); + const amount = batch.column(1); + for (let row = 0; row < batch.rowCount; row++) { + if (!price.isNull(row) && !amount.isNull(row)) { + notional += price.getDouble(row) * amount.getDouble(row); + } + } + }, + { initialCredit: 1024 * 1024 }, + ); + await query.completion; + console.log("ETH-USD notional:", notional); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +for the typed column getters. ### Compression @@ -632,45 +978,138 @@ Handle the failure at the stage where it occurs: | Local value validation | Fix the value; the row in progress was discarded. Test `QwpBatchTooLargeError` before `RangeError` because it extends `RangeError`. | | `QwpMemoryReplayAppendTimeoutError` / `QwpReplayStoreAppendTimeoutError` | The batch stays staged. Slow down and retry the flush, not the rows. | | Server rejection | See [Ingestion errors](#ingestion-errors); a terminal rejection fails the sender. | -| `QwpEgressQueryError` | Fix SQL or bind values; the lease remains usable. | +| `QwpEgressQueryError` | Check `status`: fix SQL or bind values for `PARSE_ERROR`; retry on a new lease for `CANCELLED`, such as a query cancelled by a server shutdown. The lease remains usable after a SQL error. | +| `QwpEgressSessionClosedError` | The query connection was lost with failover off. Close the lease and retry the whole query. | | `QwpReconnectExhaustedError` | Close the failed sender or query lease and borrow a new one. | ### Ingestion errors A local column or `at()` validation error discards the unfinished row. A **server** rejection can arrive after `flush()` resolves: register -`ingressSession.onSenderError` to inspect its `category`, `appliedPolicy`, -`serverStatusByte`, and `serverMessage`. Without a callback, the client logs -rejections. `waitForAcknowledged()` rejects for a rejected batch. -Branch on category, not the unstable message text, and redact messages before -sending them to external trackers: they may contain row values. Terminal +`ingressSession.onSenderError` to receive it. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + ingressSession: { + onSenderError: (error) => { + // Keep serverMessage out of external trackers: it may contain row values. + console.error("QuestDB rejected a batch", { + category: error.category, // for example "schema-mismatch" + policy: error.appliedPolicy, // "retriable", "retriable-other", "terminal", or "abandoned" + table: error.tableName, + fromFsn: error.fromFsn, + toFsn: error.toFsn, + }); + }, + }, +}); +await db.close(); +``` + +Each error has `category`, `appliedPolicy`, `serverStatusByte`, +`serverMessage`, `messageSequence`, the rejected frame range `fromFsn` to +`toFsn`, `tableName` when the server reports one, and `detectedAtMs`. +Categories and policies are lowercase, hyphenated strings +(`QWP_SENDER_ERROR_CATEGORY`, `QWP_SENDER_ERROR_POLICY`). The handler also +runs for retriable rejections, which the client resends: only `terminal` (the +sender stops) and `abandoned` (journal data set aside) mean the rows are not +being delivered. The default policy of each category is listed under +[Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). +Without a callback, the client logs rejections. `waitForAcknowledged()` +rejects for a rejected batch. Branch on category, not the unstable message +text, and redact messages before sending them to external trackers. Terminal session errors also reach `ingressSession.onError` with `terminal: true`. #### Recovering from a terminal rejection -A terminal batch rejection stops that sender. Close it and fix the row or -schema before retrying. Without `sf_dir`, unacknowledged rows on the failed -sender are lost. With `sf_dir`, the rejected batch stays at the front of the -journal and blocks **all** tables using it until fixed or deliberately -quarantined. A pooled sender is replaced after a failed `close()`, but its -replacement sees the same blocked journal. See the -[operating guide](/docs/high-availability/store-and-forward/operating-and-tuning/). +A terminal batch rejection stops that sender, and its `close()` rejects with +`QwpReplayRejectedError`. Without `sf_dir`, unacknowledged rows on the failed +sender are lost: fix the row or schema before retrying. With `sf_dir`, the +rejected batch stays at the front of the journal and blocks **all** tables +using it, across restarts. A pooled sender is replaced after a failed +`close()`, but its replacement sees the same blocked journal. To unblock the +journal: + +1. Stop the process that owns the slot: `/-` for a + pooled sender, or `/` for a standalone one. The + `onSenderError` report, or the client's log line, gives the category and + the server message. If the slot stays locked after a crash, see + [Lock recovery](#sf-lock-recovery). +2. Either fix the cause on the server so that every journaled batch is + accepted, for example by changing a conflicting column with + `ALTER TABLE ... ALTER COLUMN ... TYPE`, or move the slot directory out of + `sf_dir` to discard its unacknowledged rows. Keep the moved directory for + inspection. +3. Start the client again with the same `sf_dir` and `sender_id`. After a + server-side fix, it replays the whole journal, including the rejected + batch. After a move, it starts with an empty slot. ### Query errors SQL errors reject query iteration and `completion` with -`QwpEgressQueryError` (`status`, `message`, `requestId`). Other failures -include `QwpEgressQueryTimeoutError`, cancellation and failover exhaustion. -After cancellation, return the lease before borrowing another. Compare -`QWP_STATUS` constants instead of matching server error text. +`QwpEgressQueryError` (`status`, `message`, `requestId`); the lease remains +usable. Compare `status` with `QWP_STATUS` constants instead of matching +server error text: + +```typescript +import { connectQwpNodeClient, QWP_STATUS, QwpEgressQueryError } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query("SELECT no_such_column FROM trades"); + for await (const batch of query) console.log(batch.rowCount); + await query.completion; + } catch (error) { + if (!(error instanceof QwpEgressQueryError)) throw error; + const invalidSql = error.status === QWP_STATUS.PARSE_ERROR; + console.error(invalidSql ? "invalid SQL:" : "query failed:", error.message); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`requestId` numbers queries per connection; it is not a server-side +correlation ID. A server that shuts down cancels running queries: they fail +with `status` `QWP_STATUS.CANCELLED` instead of failing over, so retry them on +a new lease. Other failures are separate classes, not `QwpEgressQueryError`: + +- `QwpEgressQueryTimeoutError`: `timeoutMs` expired. +- `QwpEgressQueryAbandonedError`: the loop ended early; see + [Cancellation and timeouts](#cancellation-and-timeouts). +- `QwpEgressQueryCancelTimeoutError`: a cancellation did not drain in time. +- `QwpEgressSessionClosedError`: the connection was lost with failover off. +- `QwpReconnectExhaustedError`: failover gave up. + +After a cancellation or a lost connection, return the lease before borrowing +another. ### Connection-level errors The pool wraps connection creation failures in `QwpPoolResourceError`; inspect -its `cause` for an authentication `QwpUpgradeError`, a -`QwpDurableAckUnavailableError`, `QwpRoleMismatchError`, or another failure. -`QwpFailoverError.attempts` records failed endpoints. A borrow at pool -capacity times out as `QwpPoolAcquireTimeoutError`. +its `cause`. A `QwpUpgradeError` covers any failure while opening the +WebSocket, so check its `kind`: `authentication` for an HTTP `401` or `403`, +`transport` or `timeout` when QuestDB is unreachable, and others such as +`role-rejected` and `version-mismatch`. `QwpRoleMismatchError` and +`QwpDurableAckUnavailableError` extend `QwpUpgradeError`, so test for them +first. With several endpoints, the cause is a `QwpFailoverError` whose +`attempts` records each endpoint's failure; with one endpoint, it is that +endpoint's error. A borrow at pool capacity times out as +`QwpPoolAcquireTimeoutError` after `acquire_timeout_ms` (5 seconds by +default). + +An authentication rejection never moves the client to another endpoint. It is +terminal before a sender's first successful connection. After that, senders +with `sf_dir` or a background start retry it indefinitely, keep buffering, +and emit an `attempt-failed` [connection event](#connection-events) for each +failed attempt; other senders and query connections treat it as terminal. See +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). #### Connection timeouts @@ -730,30 +1169,65 @@ try { } ``` +A query that fails over restarts from its first row; see +[Query failover](#query-failover) before accumulating results. + ### Ingestion reconnect Senders resend unacknowledged batches after a disconnect. A sender in default memory mode gives up after `reconnect_max_duration_millis` (5 minutes by -default); background memory mode or `sf_dir` retries indefinitely, subject -to queue or journal capacity. The first connection is fail-fast unless -`initial_connect_retry=on` (bounded), `initial_connect_retry=async`, or -`lazy_connect=on` (background). Setting a `reconnect_*` key implicitly -requests bounded first-connection retry unless you explicitly set +default): it fails with `QwpReconnectExhaustedError` and its unacknowledged +rows are lost. Background memory mode or `sf_dir` retries indefinitely, +subject to queue or journal capacity. For the first connection, see +[Startup and outage modes](#ingestion-modes). Setting a `reconnect_*` key +implicitly requests bounded first-connection retry unless you explicitly set `initial_connect_retry=off`. ### Query failover -A lost query connection can re-execute the query **from its first row**, even -if your loop has already consumed rows. A replay can also return zero batches. -For streaming results you cannot retract, use `failover=off` and retry the -whole operation after a transport failure. Otherwise buffer the result and -reset it using `egressSession.onReplayReset`; use a single-query client for -that callback because request IDs are per connection, not unique across -pooled leases. `batch.batchSequence === 0n` detects a nonempty replay but -not a replay returning no batches. Query failover defaults to 8 attempts, -which may end before the 30-second time budget: raise -`failover_max_attempts` for longer outages. A failed query lease must be -closed and replaced. +Query failover is on by default. A lost query connection re-executes the +query **from its first row**, even if your loop has already consumed rows, and +a replay can also return zero batches. For streaming results you cannot +retract, use `failover=off` and retry the whole operation when the query fails +with `QwpEgressSessionClosedError`. Otherwise buffer the result and reset the +buffer in `egressSession.onReplayReset`. Give that callback a client with one +query connection (`query_pool_max=1`), because request IDs are per connection, +not unique across pooled leases: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const rows: unknown[][] = []; +// One query connection, so every replay reset belongs to the query below. +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;query_pool_max=1;sender_pool_min=0;", + { egressSession: { onReplayReset: () => { rows.length = 0; } } }, +); +try { + const lease = await db.borrowQuery(); + try { + const query = await lease.query( + "SELECT timestamp, symbol, price FROM trades WHERE symbol = 'ETH-USD'", + { initialCredit: 1024 * 1024 }, + ); + for await (const batch of query) { + for (const row of batch.rows()) rows.push([...row]); + } + await query.completion; + console.log(`${rows.length} rows`); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +`batch.batchSequence === 0n` detects a nonempty replay but not a replay +returning no batches. Query failover defaults to 8 attempts, which may end +before the 30-second time budget: raise `failover_max_attempts` for longer +outages. When failover gives up, the query fails with +`QwpReconnectExhaustedError`; close the lease and borrow a new one. ### Typed reconnect policy @@ -763,16 +1237,48 @@ the typed object. Ingestion `reconnect_*` keys trigger first-connection retry, but a typed ingestion `reconnect` object does not. A query's first connection retries only if `failover=on` is explicit, a `failover_*` key is set without `failover=off`, or a typed query `reconnect` object is provided. +A typed `egressSession.reconnect` object also turns query failover back on +when the connect string says `failover=off`, so do not add one, even just for +`onEvent`, to a client that must not re-execute SQL. ### Connection events -Set `ingressSession.reconnect.onEvent` or -`egressSession.reconnect.onEvent` to observe `connected`, `reconnecting`, -`reconnected`, `failed-over`, and `durable-ack-unavailable` events. No event -marks a terminal failure: use `ingressSession.onError` with `terminal: true` -for that. Connect-string `connection_listener_inbox_capacity` configures -ingestion events only; for query events, use typed -`egressSession.connectionListenerInboxCapacity`. +Set `ingressSession.reconnect.onEvent` or `egressSession.reconnect.onEvent` to +observe connection events. Each event has a `kind`, an `attempt` number, +`timestampMs`, and, where relevant, `endpoint`, `previousEndpoint`, and +`cause`. The kinds are `connected`, `reconnecting`, `attempt-failed` (one per +failed connection attempt, with its `cause`), `reconnected`, `failed-over`, +and `durable-ack-unavailable`. Orphan drainers (`drain_orphans=on`) also +report `primary-unavailable` and `durable-ack-persistent-failure`. +`QWP_RECONNECT_EVENT_KIND` lists them all. + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + ingressSession: { + // A typed reconnect object replaces the connect-string policy: set its limits too. + reconnect: { + maxDurationMs: 300_000, + onEvent: (event) => { + if (event.kind === "attempt-failed") { + console.warn("connection attempt failed", event.endpoint, event.cause); + } else { + console.info("connection event", event.kind, event.endpoint); + } + }, + }, + }, +}); +await db.close(); +``` + +No event marks a terminal failure: use `ingressSession.onError` with +`terminal: true` for that. An `egressSession.reconnect` object added for +`onEvent` re-enables query failover on a `failover=off` client; see +[Typed reconnect policy](#typed-reconnect-policy). Connect-string +`connection_listener_inbox_capacity` configures ingestion events only; for +query events, use typed `egressSession.connectionListenerInboxCapacity`. ## Concurrency @@ -785,24 +1291,57 @@ distinct `sender_id` values. ## Configuration reference The [connect string reference](/docs/connect/clients/connect-string/) lists -shared keys and defaults. Important Node.js differences: +shared keys and defaults; Node.js differences follow the typed options. + +### Programmatic options + +The second argument of `connectQwpNodeClient` accepts typed `sender`, +`ingressSession`, `egressSession`, `egress`, `webSocket`, `storeAndForward`, and +`pool` options. For example: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { + sender: { awaitServerAck: true }, + ingressSession: { + onSenderError: (error) => + console.error("rejected batch:", error.category, error.serverMessage), + }, + pool: { senderPoolMax: 2 }, +}); +await db.close(); +``` + +Typed options override the corresponding connect-string settings. A typed +`reconnect` object replaces the whole reconnect policy from the string; see +[Typed reconnect policy](#typed-reconnect-policy). For the full typed API, see +the +[client reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). ### Differences from other clients | Area | Node.js behavior | |---|---| -| Startup/outage | `lazy_connect=on` starts senders in the background; default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. | +| Startup/outage | `lazy_connect=on` starts senders in the background; default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. See [Startup and outage modes](#ingestion-modes). | +| Authentication | After a sender's first successful connection, `401`/`403` is retried indefinitely by senders with `sf_dir` or a background start; other senders and queries treat it as terminal. | | Query startup | First-connect retry requires explicit `failover=on`, a `failover_*` key (without `failover=off`), or typed `egressSession.reconnect`. | | `target`, `zone` | Also apply to ingestion when set in the connect string; use typed `egress.target` for queries only. | -| `sf_dir` | Node.js recursively creates missing parents and a slot. Its `.lock.owner` directory can outlive a crashed process; see [Lock recovery](#sf-lock-recovery). | +| `sf_dir` | Node.js recursively creates missing parents and a slot; pooled senders use slots named `-`. Its `.lock.owner` directory can outlive a crashed process; see [Lock recovery](#sf-lock-recovery). | | SF-only keys | Explicit `sf_durability` (even `memory`), `sf_sync_interval_millis`, `drain_orphans`, `max_background_drainers`, and `catch_up_cap_gap_min_escalation_window_millis` require `sf_dir`. `sf_durability=append` is supported. | | `sf_max_total_bytes` | With `sf_dir` it is a journal size **target**, not a disk quota; without `sf_dir` it caps the memory queue. | | Durable ACK | Background-started senders, including with `sf_dir`, retry an unavailable capability from startup. Explicit `durable_ack_keepalive_interval_millis` also requests durable ACK even at `0`; negatives are rejected. | +| Timeouts | `connect_timeout` defaults to 15 seconds and also bounds DNS and the TLS handshake; `auth_timeout_ms` defaults to `connect_timeout`. `close_flush_timeout_millis` defaults to 5 seconds. When it expires, a standalone sender's `close()` rejects with `QwpSenderCloseTimeoutError`, while `db.close()` resolves and usually reports a non-terminal `QwpIngressAckTimeoutError` to `ingressSession.onError`. | +| `poison_min_escalation_window_millis` | Defaults to 300000 (5 minutes), not 5000. | +| Auto-flush | `auto_flush_interval` counts from the last flush (or sender creation), not from the first buffered row. `auto_flush_bytes` defaults to 0 (off). | +| `max_lifetime_ms` | Closes only idle connections above the pool minimum; the minimum connections are never recycled. | | `connection_listener_inbox_capacity` | Configures only ingestion events, not query events. | -| Parsing | Use `0`, not `off`, for `auto_flush_rows` and `auto_flush_interval`; sizes take single-letter suffixes (`k`, `m`, `g`, `t`). `compression_level` requires `compression=zstd` or `auto`. `tls_roots` must be PEM; `tls_roots_password` is rejected. | +| Parsing | Use `0`, not `off`, for `auto_flush_rows` and `auto_flush_interval`. Size keys take single-letter suffixes (`k`, `m`, `g`, `t`), but `initial_credit` takes plain bytes. `compression_level` requires `compression=zstd` or `auto`. `tls_roots` must be PEM, and `tls_roots_password` is rejected; without `tls_roots`, certificates are checked against Node.js's bundled CA store. `init_buf_size` and `max_buf_size` are rejected on `ws` and `wss` strings. | | Reconnect jitter | Ingestion uses full jitter, unlike the shared guide's equal-jitter ingestion schedule. | -| Typed SF options | A typed-only `storeAndForward` object defaults to append durability, 1 GiB, and fail-on-full; connect strings default to memory durability, 10 GiB, and wait-on-full. | -| Standalone `Sender` | Ignores pool and query-only keys with a warning; `on_*_error` keys are accepted but not applied. | +| Error categories | `onSenderError` reports categories and policies as lowercase, hyphenated strings, such as `schema-mismatch` and `retriable-other`, and adds the `cancelled` and `limit-exceeded` categories. | +| `on_*_error` | Accepted by every entry point but not applied; observe rejections with `ingressSession.onSenderError`. | +| Typed SF options | A typed `storeAndForward` option passed with a connect string, as in `connectQwpNodeClient(conf, { storeAndForward })`, keeps the connect-string defaults: memory durability, 10 GiB, and wait-on-full. Without a connect string, as in `connectQwpNodeSender({ url, storeAndForward })`, the defaults are append durability, 1 GiB, and fail-on-full. | +| Standalone `Sender` | Ignores pool and query-only keys with a warning, but applies `client_id` and `lazy_connect`. | ## Migration @@ -813,7 +1352,9 @@ or `wss::`, followed by `await sender.connect()`. QWP adds querying, replay, store-and-forward, and richer types. Unlike ILP HTTP, a QWP `flush()` can resolve **before** the server accepts the rows, so wait for an ACK when committing source offsets. Legacy ILP-only keys such as `retry_timeout`, -`init_buf_size`, and `tls_ca` are rejected on QWP strings. +`init_buf_size`, and `tls_ca` are rejected on QWP strings. The +[Standalone Sender](#standalone-sender) example shows the QWP form of the row +API; [ILP transports](#ilp-transports-legacy) shows the ILP form. ### Upgrading from 4.x @@ -825,11 +1366,36 @@ transports. ## ILP transports (legacy) -Use `Sender.fromConfig("http::addr=localhost:9000;")` for ILP over HTTP, -or `tcp::addr=localhost:9009;` for TCP. ILP is ingestion-only: call -`sender.flush()` before `sender.close()` or buffered rows are lost. See the -[ILP overview](/docs/connect/compatibility/ilp/overview/) for authentication, -protocol versions, and transport configuration. +The standalone `Sender` still speaks InfluxDB Line Protocol (ILP) over HTTP +and TCP, for ingestion only. ILP over HTTP sends each `flush()` as an HTTP +request and has no `connect()` step; calling it throws: + +```typescript +import { Sender } from "@questdb/nodejs-client"; + +const sender = await Sender.fromConfig("http::addr=localhost:9000;"); +try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "sell") + .floatColumn("price", 2615.54) + .floatColumn("amount", 0.00044) + .at(Date.now(), "ms"); + await sender.flush(); // close() does not flush: unflushed rows are lost +} finally { + await sender.close(); +} +``` + +For authentication, add `username=...;password=...;` (or, on QuestDB +Enterprise, `token=...;`) to the connect string. For ILP over TCP, use +`tcp::addr=localhost:9009;` and call `await sender.connect()` before writing. +`Sender.fromEnv()` reads the connect string from the `QDB_CLIENT_CONF` +environment variable. ILP uses the same nine column methods as the standalone +QWP `Sender`; see [Column methods](#column-methods). See the +[ILP overview](/docs/connect/compatibility/ilp/overview/) for TCP +authentication, protocol versions, and transport configuration. ## Next steps @@ -837,3 +1403,5 @@ protocol versions, and transport configuration. - [Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/). - [Connect string reference](/docs/connect/clients/connect-string/) and the [Node.js API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html). +- [`@questdb/nodejs-client` on npm](https://www.npmjs.com/package/@questdb/nodejs-client) + and its [source on GitHub](https://github.com/questdb/nodejs-questdb-client). diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index 0a2bb2db11..d4a2fc0cc9 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -157,7 +157,8 @@ length, and what bounds your tolerance is buffer capacity `QwpReconnectExhaustedError`; see the [Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). - Jitter: **equal-jitter** `[base, 2·base)` — non-zero lower bound damps - reconnect storms when many producers share a cluster + reconnect storms when many producers share a cluster. The Node.js client + uses full jitter, `[0, base)`, instead - Inter-host pause within a round: **none** — the client walks the full address list as fast as `auth_timeout_ms` allows, paying one backoff sleep at round exhaustion @@ -198,7 +199,7 @@ will not help. | Condition | Why terminal | |---|---| | HTTP `401` / `403` on upgrade | Credentials are assumed to be cluster-wide. Some clients retry after a sender's first successful connection; see [Authentication is cluster-wide](#authentication-is-cluster-wide). | -| Server-status reject (SF) | Application-layer reject; replay reproduces the same response. | +| Server rejection with a terminal policy: by default `SCHEMA_MISMATCH`, `PARSE_ERROR`, `SECURITY_ERROR`, and `PROTOCOL_VIOLATION`, or a batch that keeps being rejected | Replaying the same bytes reproduces the rejection. Retriable categories, such as `WRITE_ERROR`, are resent instead; see [Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). | ### Topology — handled inside the round diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 0d751ea8c0..638aeeeea0 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -52,7 +52,7 @@ for the full list. The failover-relevant keys are: |---|---|---|---| | `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely, so raising this does nothing for failover windows. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Starting backoff sleep at round exhaustion. Doubles up to `reconnect_max_backoff_millis`. | -| `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. | +| `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter, the actual sleep lands in `[max, 2·max)` once the base saturates. The Node.js client uses full jitter, so its sleep lands in `[0, max)`. | | `initial_connect_retry` | `off` \| `on` \| `async` | `off` | Whether to apply the same retry loop to the very first connect attempt. See below. | ### `initial_connect_retry` diff --git a/documentation/high-availability/store-and-forward/configuration.md b/documentation/high-availability/store-and-forward/configuration.md index cacfd3bf6f..8c58e865b3 100644 --- a/documentation/high-availability/store-and-forward/configuration.md +++ b/documentation/high-availability/store-and-forward/configuration.md @@ -50,7 +50,7 @@ and host-walk semantics are documented in |---|---|---|---| | `reconnect_max_duration_millis` | int (ms) | `300000` (5 min) | Bounds the blocking sync initial connect only (`initial_connect_retry=on`/`sync`). A running sender's reconnect loop never consults it and retries indefinitely. Exception: a Node.js sender in default memory mode, without `sf_dir`, `initial_connect_retry=async`, or `lazy_connect=on`, applies it to every outage; see [the Node.js client](/docs/connect/clients/nodejs/#ingestion-reconnect). | | `reconnect_initial_backoff_millis` | int (ms) | `100` | Initial backoff sleep at round exhaustion. | -| `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`. | +| `reconnect_max_backoff_millis` | int (ms) | `5000` | Cap on the exponential backoff. With equal-jitter the actual sleep lands in `[max, 2·max)`; the Node.js client uses full jitter, `[0, max)`. | | `initial_connect_retry` | enum | `off` | `off` (alias `false`): first-connect failure is terminal. `on` (aliases `sync`, `true`): same retry loop as reconnect, blocking the constructor. `async`: same retry loop in the I/O thread, non-blocking. | | `close_flush_timeout_millis` | int (ms) | `60000` on Java and .NET; `5000` on Rust, C, C++, Python, Go and Node.js | `close()` blocks up to this long waiting for `ackedFsn ≥ publishedFsn`. `0` or `-1` skips the drain wait. The safety-net `checkError()` still runs. | From 32056fc65367ed4659ac294f22294ee9d385fd45 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 14:17:58 +0100 Subject: [PATCH 21/25] docs(nodejs): fix startup, batch size, transaction, and failover guidance - Startup: initial_connect_retry=on and reconnect_* keys retry senders only. connectQwpNodeClient() still rejects at once with the default query_pool_min=1; document query_pool_min=0, and lazy_connect=on as initial_connect_retry=async plus query_pool_min=0. Add Node.js notes to the connect-string startup recipe and the lazy_connect entry, and correct the recipe's claim that the budget covers later reconnects. - Batch size: a frame built offline after any background start can exceed the server limit and block replay, with or without sf_dir; recommend sf_max_segment_bytes=1m for every background start. - Transactions: a pooled sender cannot roll back, because close() commits batches that auto-flush already sent. Show the standalone Sender, whose close() rolls back an uncommitted transaction. - Zero-copy views: the example accumulates a total, so set failover=off to stop a re-run from adding replayed batches twice. --- .../connect/clients/connect-string.md | 24 ++++- documentation/connect/clients/nodejs.md | 96 +++++++++++-------- 2 files changed, 76 insertions(+), 44 deletions(-) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index f7c9e3fc7e..c4a145acac 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -171,11 +171,22 @@ set the role with the typed `egress` option instead, as described under ws::addr=node-a:9000;reconnect_max_duration_millis=120000; ``` -The 2-minute reconnect budget covers both the *first* connect and any -subsequent reconnect: setting any explicit `reconnect_*` key implicitly -turns on `initial_connect_retry`. See +Setting any explicit `reconnect_*` key implicitly turns on +`initial_connect_retry`, so the sender retries its *first* connect for up to +2 minutes. A running sender retries later outages indefinitely. See [Ingress reconnect](#reconnect-keys). +:::caution Node.js client + +On the Node.js client, the retry covers senders only. +`connectQwpNodeClient()` also opens a query connection, which still gives up +almost at once, so add `query_pool_min=0`, or use `lazy_connect=on` to start +without waiting. In default memory mode, the budget also ends every later +outage after 2 minutes. See +[Node.js startup and outage modes](/docs/connect/clients/nodejs/#ingestion-modes). + +::: + ## Recipes {#recipes} Goal-to-keys mapping. For complete connect-string templates, see @@ -857,7 +868,12 @@ explicit setter always wins over the string. - `lazy_connect` — when `on`, the pool defers opening its first connection until the first borrow, so construction succeeds against a server that is down. This is the supported way to tolerate a server that starts after your - application. Default: `off`. + application. Default: `off`. The Node.js client instead starts its senders + connecting in the background (`initial_connect_retry=async`) and sets + `query_pool_min=0`, rejecting a positive value. Those senders also retry an + outage indefinitely rather than giving up after + `reconnect_max_duration_millis`; see + [Node.js startup and outage modes](/docs/connect/clients/nodejs/#ingestion-modes). ## Error handling {#error-handling} diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 2949a3d3c2..e5a7d9d6b9 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -183,13 +183,17 @@ borrows, and close the client on shutdown. Two independent choices decide what a sender does while QuestDB is unreachable: -- **Startup.** A *foreground start*, the default, connects inside +- **Startup.** A *foreground start*, the default, connects senders inside `connectQwpNodeClient()`, which rejects if QuestDB is down. A *background - start* (`lazy_connect=on`) returns immediately and connects senders in the - background. `lazy_connect=on` sets `query_pool_min` to 0 and rejects a - positive value, and a query borrowed before QuestDB is reachable fails. A - standalone `Sender` gets a background start with - `initial_connect_retry=async`. + start* connects senders in the background: `initial_connect_retry=async` + selects it, and `lazy_connect=on` selects it and also sets + `query_pool_min=0`, rejecting a positive value. `connectQwpNodeClient()` + returns while QuestDB is down only with `query_pool_min=0`: with + `initial_connect_retry=async` alone, the default query connection still + has to connect at startup. A query borrowed before QuestDB is reachable + fails. A standalone `Sender` gets a background start from either key. With + a background start, also set `sf_max_segment_bytes=1m`; see + [Batch size limits](#batch-size-limits). - **Storage.** Without `sf_dir`, unacknowledged rows live in memory and are lost if the process exits. With `sf_dir`, they are journaled to disk and replayed after a restart; see [Store-and-forward](#store-and-forward). @@ -197,13 +201,16 @@ unreachable: | Mode | Enabled by | QuestDB down at startup | During an outage | |---|---|---|---| | Default memory | Neither `sf_dir` nor a background start | Startup rejects | `flush()`, and `at()` when it triggers an auto-flush, wait for the reconnect for up to `reconnect_max_duration_millis` (5 minutes by default). Then the sender fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost | -| Background memory | A background start without `sf_dir` | Starts; rows queue in memory | Rows queue in memory, up to `sf_max_total_bytes` (128 MiB by default); retries continue indefinitely | -| Store-and-forward | `sf_dir`, with either startup | Foreground: startup rejects. Background: starts | Rows go to the disk journal; retries continue indefinitely, from startup with a background start or after the first successful connection otherwise | +| Background memory | A background start without `sf_dir` | Starts if `query_pool_min=0`, as with `lazy_connect=on`; rows queue in memory | Rows queue in memory, up to `sf_max_total_bytes` (128 MiB by default); retries continue indefinitely | +| Store-and-forward | `sf_dir`, with either startup | Foreground: startup rejects. Background: starts if `query_pool_min=0` | Rows go to the disk journal; retries continue indefinitely, from startup with a background start or after the first successful connection otherwise | -A foreground start fails fast unless `initial_connect_retry=on` or any -`reconnect_*` key is set; then it retries the first connection for up to -`reconnect_max_duration_millis` before rejecting. A locked journal fails -startup even with a background start. +A foreground start fails fast. `initial_connect_retry=on`, or any +`reconnect_*` key, makes senders retry their first connection for up to +`reconnect_max_duration_millis` before rejecting. These keys do not apply to +query connections, so with the default `query_pool_min=1`, +`connectQwpNodeClient()` still rejects almost at once: also set +`query_pool_min=0` to wait for the senders. A locked journal fails startup +even with a background start. ### Closing the pooled client @@ -374,10 +381,11 @@ QuestDB advertises its maximum batch size on connection (about 2 MiB on a default server). The client splits a larger batch into several frames at row boundaries. A single row larger than the limit fails with `QwpBatchTooLargeError`: call `reset()` and shrink that row, for example a -large VARCHAR or BINARY value. A sender that starts offline cannot yet know -the server limit. With `sf_dir` and `lazy_connect=on`, set -`sf_max_segment_bytes=1m`: it also caps each journaled frame at 1 MiB, so an -oversized frame cannot block journal replay. +large VARCHAR or BINARY value. A sender with a background start cannot know +the limit before its first connection, so a frame built while QuestDB is +down can exceed it. That frame is then never delivered: it is retried +indefinitely and blocks every later batch. With a background start, with or +without `sf_dir`, set `sf_max_segment_bytes=1m` to cap each frame at 1 MiB. ### Awaiting acknowledgements @@ -453,31 +461,33 @@ Set `transaction=on` to defer server commits of auto-flushed batches until `flush()` (or `commit()` on a pooled sender). Transactions are atomic per table, not across tables, and QuestDB can commit early when the table exceeds [`qwp.max.uncommitted.rows`](/docs/configuration/qwp/#qwpmaxuncommittedrows). -Closing a standalone sender without `flush()` rolls back the open transaction; -returning a pooled sender with `close()` flushes and commits instead. + +Only a standalone sender can roll back: closing it without `flush()` discards +the open transaction. A pooled sender has no rollback. Returning it with +`close()` flushes and commits the open transaction, including batches that +auto-flush already sent; `reset()` drops only rows that have not been sent. +If your code throws partway through a batch, a pooled sender therefore +commits the rows written before the error. When a failed batch must leave no +rows, use a standalone sender: ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { Sender } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;transaction=on;"); +const sender = await Sender.fromConfig("ws::addr=localhost:9000;transaction=on;"); try { - const sender = await db.borrowSender(); - try { - for (const [side, price] of [["buy", 2615.54], ["sell", 2615.62]] as const) { - await sender - .table("trades") - .symbol("symbol", "ETH-USD") - .symbol("side", side) - .doubleColumn("price", price) - .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); - } - await sender.commit(); // both rows become visible together - } finally { - await sender.close(); + await sender.connect(); + for (const [side, price] of [["buy", 2615.54], ["sell", 2615.62]] as const) { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", side) + .floatColumn("price", price) // DOUBLE: the Sender has no doubleColumn() + .floatColumn("amount", 0.5) + .at(Date.now(), "ms"); } + await sender.flush(); // commits: both rows become visible together } finally { - await db.close(); + await sender.close(); // rolls back the transaction if flush() was not reached } ``` @@ -932,7 +942,8 @@ them. ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; -const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +// failover=off: a re-run after a lost connection would count batches twice. +const db = await connectQwpNodeClient("ws::addr=localhost:9000;failover=off;"); try { const lease = await db.borrowQuery(); try { @@ -960,7 +971,11 @@ try { } ``` -See the [client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) +The callback adds to a running total, so the example turns failover off: with +failover on, a lost connection re-runs the query and its batches are added +again. If the query fails with `QwpEgressSessionClosedError`, run it again +from an empty total; see [Query failover](#query-failover). See the +[client API reference](https://questdb.github.io/nodejs-questdb-client/modules/_questdb_nodejs-client.html) for the typed column getters. ### Compression @@ -1180,8 +1195,9 @@ default): it fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost. Background memory mode or `sf_dir` retries indefinitely, subject to queue or journal capacity. For the first connection, see [Startup and outage modes](#ingestion-modes). Setting a `reconnect_*` key -implicitly requests bounded first-connection retry unless you explicitly set -`initial_connect_retry=off`. +implicitly requests bounded first-connection retry for senders unless you +explicitly set `initial_connect_retry=off`; `connectQwpNodeClient()` waits +for that retry only with `query_pool_min=0`. ### Query failover @@ -1323,7 +1339,7 @@ the | Area | Node.js behavior | |---|---| -| Startup/outage | `lazy_connect=on` starts senders in the background; default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. See [Startup and outage modes](#ingestion-modes). | +| Startup/outage | `lazy_connect=on` or `initial_connect_retry=async` starts senders in the background, but `connectQwpNodeClient()` starts while QuestDB is down only with `query_pool_min=0`, which `lazy_connect=on` sets. First-connection retry from `initial_connect_retry=on` or a `reconnect_*` key covers senders only. Default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. See [Startup and outage modes](#ingestion-modes). | | Authentication | After a sender's first successful connection, `401`/`403` is retried indefinitely by senders with `sf_dir` or a background start; other senders and queries treat it as terminal. | | Query startup | First-connect retry requires explicit `failover=on`, a `failover_*` key (without `failover=off`), or typed `egressSession.reconnect`. | | `target`, `zone` | Also apply to ingestion when set in the connect string; use typed `egress.target` for queries only. | From 775becf6bb0f6a1a4b976359b9cf237bd377c054 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 16:34:44 +0100 Subject: [PATCH 22/25] docs(nodejs): fix quick start, ACK timeouts, and recovery guidance - Quick start polls for the row it wrote, so a first run on a fresh table prints it instead of nothing - Document ackTimeoutMs and durableAckTimeoutMs: a timed-out flush() rejects with a plain Error, waitForAcknowledged() with QwpIngressAckTimeoutError, and neither means the batch failed; raise the durable-ACK deadline above the replication throttle window - Explain that tables needing DEDUP must exist before the first ingester runs, and show ALTER TABLE ... DEDUP ENABLE for an auto-created table - Make Startup and outage modes the single source: recommended keys first, a mode table, a caution that default memory mode blocks producers, and bullets instead of conditional chains; cut restatements elsewhere - Add a runnable decimals example and a decimal bind, and move the old decimal anchors back into the section - Add Quarantined journal slots: .unreplayable-N and .failed slots, and replay with retryQwpNodeOrphanSlot() - Show the onError event, state that every named class is exported, that same-host stale locks are reclaimed automatically, and that at() writes an existing table's designated timestamp - Align connect-string, client-failover, and store-and-forward pages with the Node.js startup and role-filter behavior --- .../connect/clients/connect-string.md | 4 +- documentation/connect/clients/nodejs.md | 362 +++++++++++++----- .../client-failover/concepts.md | 2 +- .../client-failover/configuration.md | 2 +- .../store-and-forward/operating-and-tuning.md | 9 +- 5 files changed, 276 insertions(+), 103 deletions(-) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index c4a145acac..3c7a7adbc1 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -844,8 +844,8 @@ per-language names. Every client now leads with a pooled facade, so these keys are a first-contact concern. The `Sender` and query-client parsers accept and ignore them; the facade reads them off the string. The Node.js `Sender` logs a warning for the -pool keys it ignores, and applies `lazy_connect`, which starts it in -background memory mode. Each has an equivalent builder setter, and an +pool keys it ignores, and applies `lazy_connect`, which gives it a +[background start](/docs/connect/clients/nodejs/#ingestion-modes). Each has an equivalent builder setter, and an explicit setter always wins over the string. - `sender_pool_min` — senders kept open even when idle. `0` lets the pool close diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index e5a7d9d6b9..d985e3e275 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -33,9 +33,10 @@ and remove TypeScript annotations. ## Quick start -Create a table, publish a row, wait for QuestDB's acknowledgement, and query -it. Acknowledged rows are applied asynchronously, so the query may initially -return no rows; see [Read-after-write](#read-after-write). +Create a table, publish a row, wait for QuestDB's acknowledgement, then poll +until the row is visible. An acknowledgement means QuestDB has committed the +row, but queries see it only after it is applied, which happens +asynchronously; see [Read-after-write](#read-after-write). ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -52,6 +53,7 @@ try { ); await ddl.completion; + const eventTime = Date.now(); // milliseconds; also used to find the row const sender = await db.borrowSender(); try { await sender @@ -60,20 +62,36 @@ try { .symbol("side", "buy") .doubleColumn("price", 2615.54) .doubleColumn("amount", 0.5) - .at(Date.now(), "ms"); + .at(eventTime, "ms"); await sender.flush(); - await sender.waitForAcknowledged(sender.publishedSequence); + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); } finally { await sender.close(); // return it to the pool } - const query = await lease.query( - "SELECT timestamp, symbol, price FROM trades LIMIT 10", - ); - for await (const batch of query) { - for (const row of batch.rows()) console.log(row); + // Poll until the acknowledged row is visible, for up to 10 seconds. + const deadline = Date.now() + 10_000; + let found = false; + while (!found && Date.now() < deadline) { + const query = await lease.query( + "SELECT timestamp, symbol, side, price, amount FROM trades " + + "WHERE timestamp = $1", + { + binds: (binds) => + binds.setTimestampMicros(0, BigInt(eventTime) * 1000n), + }, + ); + for await (const batch of query) { + for (const row of batch.rows()) { + // [timestamp in microseconds, symbol, side, price, amount] + console.log(row); + found = true; + } + } + await query.completion; + if (!found) await new Promise((resolve) => setTimeout(resolve, 100)); } - await query.completion; + if (!found) throw new Error("row not visible after 10 seconds"); } finally { await lease.close(); } @@ -180,37 +198,62 @@ borrows, and close the client on shutdown. ### Startup and outage modes {#ingestion-modes} -Two independent choices decide what a sender does while QuestDB is -unreachable: - -- **Startup.** A *foreground start*, the default, connects senders inside - `connectQwpNodeClient()`, which rejects if QuestDB is down. A *background - start* connects senders in the background: `initial_connect_retry=async` - selects it, and `lazy_connect=on` selects it and also sets - `query_pool_min=0`, rejecting a positive value. `connectQwpNodeClient()` - returns while QuestDB is down only with `query_pool_min=0`: with - `initial_connect_retry=async` alone, the default query connection still - has to connect at startup. A query borrowed before QuestDB is reachable - fails. A standalone `Sender` gets a background start from either key. With - a background start, also set `sf_max_segment_bytes=1m`; see - [Batch size limits](#batch-size-limits). +To start while QuestDB is down and keep accepting rows during an outage, add +`lazy_connect=on;sf_max_segment_bytes=1m;` to the connect string. To also +keep unacknowledged rows across process restarts, add +`sf_dir=/var/lib/myapp/qdb-sf;sender_id=trades;`. These keys select one of +three modes: + +| Mode | Keys | QuestDB down at startup | During an outage | +|---|---|---|---| +| Default memory | None | `connectQwpNodeClient()` rejects | `flush()`, and `at()` when it triggers an auto-flush, wait for the reconnect for up to `reconnect_max_duration_millis` (5 minutes by default). Then the sender fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost | +| Background memory | `lazy_connect=on` | Starts; rows queue in memory | Rows queue in memory, up to `sf_max_total_bytes` (128 MiB by default); retries continue indefinitely | +| Store-and-forward | `sf_dir`, usually with `lazy_connect=on` | With `lazy_connect=on`, starts and journals rows; without it, rejects | Rows go to the disk journal and survive a process restart; retries continue indefinitely | + +:::caution Default memory mode blocks producers + +Auto-flush runs on the first `at()` call 100 ms or more after the last flush, +so in default memory mode a producer blocks almost as soon as an outage +starts. The blocked call throws if QuestDB is still unreachable after +`reconnect_max_duration_millis` (5 minutes by default). If a producer must +not block on QuestDB, for example an HTTP request handler, use a background +start: its calls block only when the replay queue is full; see +[Backpressure](#backpressure). + +::: + +How the keys combine: + +- **Foreground and background start.** By default, senders get a + *foreground start*: `connectQwpNodeClient()` connects them and rejects if + QuestDB is down. `lazy_connect=on` gives them a *background start* instead + and sets `query_pool_min=0`; combining it with a positive + `query_pool_min` is a configuration error. `initial_connect_retry=async` + also gives senders a background start, but `connectQwpNodeClient()` still + opens one query connection at startup and rejects while QuestDB is down, + unless you also set `query_pool_min=0`. A query borrowed before QuestDB is + reachable fails. A standalone `Sender` gets a background start from either + key. +- **Batch size.** With a background start, set `sf_max_segment_bytes=1m`, + with or without `sf_dir`; see [Batch size limits](#batch-size-limits). - **Storage.** Without `sf_dir`, unacknowledged rows live in memory and are lost if the process exits. With `sf_dir`, they are journaled to disk and replayed after a restart; see [Store-and-forward](#store-and-forward). - -| Mode | Enabled by | QuestDB down at startup | During an outage | -|---|---|---|---| -| Default memory | Neither `sf_dir` nor a background start | Startup rejects | `flush()`, and `at()` when it triggers an auto-flush, wait for the reconnect for up to `reconnect_max_duration_millis` (5 minutes by default). Then the sender fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost | -| Background memory | A background start without `sf_dir` | Starts if `query_pool_min=0`, as with `lazy_connect=on`; rows queue in memory | Rows queue in memory, up to `sf_max_total_bytes` (128 MiB by default); retries continue indefinitely | -| Store-and-forward | `sf_dir`, with either startup | Foreground: startup rejects. Background: starts if `query_pool_min=0` | Rows go to the disk journal; retries continue indefinitely, from startup with a background start or after the first successful connection otherwise | - -A foreground start fails fast. `initial_connect_retry=on`, or any -`reconnect_*` key, makes senders retry their first connection for up to -`reconnect_max_duration_millis` before rejecting. These keys do not apply to -query connections, so with the default `query_pool_min=1`, -`connectQwpNodeClient()` still rejects almost at once: also set -`query_pool_min=0` to wait for the senders. A locked journal fails startup -even with a background start. + `sf_dir` alone keeps a foreground start: startup rejects while QuestDB is + down, and outages after the first successful connection are retried + indefinitely. +- **Tables.** An ingester that starts while QuestDB is down cannot create + its tables first. Create tables that need DEDUP beforehand; see + [Store-and-forward](#store-and-forward). +- **First-connection retry.** `initial_connect_retry=on`, or any + `reconnect_*` key unless `initial_connect_retry=off` is set, makes senders + retry their first connection for up to `reconnect_max_duration_millis` + instead of failing at once. The sender stays in default memory mode. These + keys do not apply to query connections, so with the default + `query_pool_min=1`, `connectQwpNodeClient()` still rejects almost at once: + also set `query_pool_min=0`. +- **Locked journal.** A journal locked by another process fails startup, + even with a background start; see [Lock recovery](#sf-lock-recovery). ### Closing the pooled client @@ -276,7 +319,9 @@ without the unit the row lands in 1970. Nanoseconds require a `bigint`. `atNow()` asks QuestDB to assign arrival time, which changes on replay. Use the event's timestamp for deduplication. For a newly created table, `"ns"` creates a TIMESTAMP_NS designated timestamp; the other units create -TIMESTAMP. The default designated column name is `timestamp`. +TIMESTAMP. The default designated column name is `timestamp`. On an existing +table, `at()` writes the table's designated timestamp column, whatever its +name. ### Null values @@ -291,18 +336,58 @@ the last flush. ### Decimals -Use `decimalColumnText(name, "0.0750")` to preserve the input scale, including -trailing zeros. Both strings and numbers accept exponents such as -`"1.5e-3"`; a JavaScript number cannot retain trailing zeros. Binary methods -`decimal64Column(name, unscaled, scale)`, `decimal128Column()`, and -`decimal256Column()` take an unscaled `bigint`. Pre-create a table if you need -a specific precision: QWP auto-creation chooses the maximum precision for the -wire width. The server currently cannot return DECIMAL with precision 9 or -less over QWP; cast it to a wider precision when querying. +Pre-create a table if you need a specific precision: QWP auto-creation +chooses the maximum precision for the wire width. + +```questdb-sql +CREATE TABLE IF NOT EXISTS trade_fees ( + timestamp TIMESTAMP, + symbol SYMBOL, + settled_price DECIMAL(18, 2), + commission DECIMAL(18, 4) +) TIMESTAMP(timestamp) PARTITION BY DAY; +``` + +`decimalColumnText(name, value)` sends a decimal as text and preserves its +scale, including trailing zeros. Both strings and numbers accept exponents +such as `"1.5e-3"`; a JavaScript number cannot retain trailing zeros. + +The binary methods `decimal64Column(name, unscaled, scale)`, +`decimal128Column()`, and `decimal256Column()` take the unscaled value as a +`bigint`, followed by the scale: + +```typescript +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); +try { + const sender = await db.borrowSender(); + try { + await sender + .table("trade_fees") + .symbol("symbol", "ETH-USD") + .decimalColumnText("settled_price", "2615.50") // keeps the trailing zero + .decimal64Column("commission", -750n, 4) // -0.0750 + .at(Date.now(), "ms"); + await sender.flush(); + } finally { + await sender.close(); + } +} finally { + await db.close(); +} +``` + +Bind parameters take the scale **before** the unscaled value, the reverse of +`decimal64Column()`: `binds.setDecimal64(0, 2, 261550n)` binds `2615.50`; see +[Bind parameters](#bind-parameters). Query results return a DECIMAL as +`{ unscaled, scale }`. The server currently cannot return DECIMAL with +precision 9 or less over QWP; cast it to a wider precision when querying. + ### Arrays `arrayColumn(name, value)` sends a uniformly shaped nested array of numbers @@ -381,7 +466,7 @@ QuestDB advertises its maximum batch size on connection (about 2 MiB on a default server). The client splits a larger batch into several frames at row boundaries. A single row larger than the limit fails with `QwpBatchTooLargeError`: call `reset()` and shrink that row, for example a -large VARCHAR or BINARY value. A sender with a background start cannot know +large VARCHAR or BINARY value. A sender with a [background start](#ingestion-modes) cannot know the limit before its first connection, so a frame built while QuestDB is down can exceed it. That frame is then never delivered: it is retried indefinitely and blocks every later batch. With a background start, with or @@ -393,12 +478,21 @@ After `flush()`, wait for the cumulative watermark: `await sender.waitForAcknowledged(sender.publishedSequence, 10_000)`. `publishedSequence` includes batches sent by auto-flush; `acknowledgedSequence` -is the last accepted one. `waitForAcknowledged()` rejects on timeout or server -rejection. A timeout alone does not mean the batch was rejected: it may still -be in flight. **Do not use the return value of `flushAndGetSequence()` as the -watermark for all your rows**: it returns `-1n` if an earlier auto-flush -already published them. To make each flush wait, use typed -`sender: { awaitServerAck: true }`. +is the last accepted one. `waitForAcknowledged()` rejects on a server +rejection, or with `QwpIngressAckTimeoutError` on timeout; without a timeout +argument, it waits up to `ackTimeoutMs` (15 seconds by default). A timeout +alone does not mean the batch was rejected: it may still be in flight. **Do +not use the return value of `flushAndGetSequence()` as the watermark for all +your rows**: it returns `-1n` if an earlier auto-flush already published them. + +To make each `flush()` wait for the acknowledgement, pass typed +`sender: { awaitServerAck: true }` as the second argument of +`connectQwpNodeClient()`; see [Programmatic options](#programmatic-options). +Each such flush waits up to `ackTimeoutMs`, then rejects with a plain `Error` +whose message starts with `timed out waiting for QWP ACK`. As with +`waitForAcknowledged()`, QuestDB may still acknowledge the batch later. Set the +deadline with typed `ingressSession: { ackTimeoutMs }`; it has no +connect-string key. #### Committing source offsets @@ -453,7 +547,9 @@ try { ``` Request [durable acknowledgement](#durable-acknowledgement) if the offset must -also survive a primary failure. +also survive a primary failure. The watermark then advances only after the +upload to object storage, so allow for the upload interval in your +`waitForAcknowledged()` timeout. ### Transactions @@ -495,10 +591,19 @@ try { Set `sf_dir` to journal batches across process restarts. To start while QuestDB is down, add a background start (`lazy_connect=on`), as below; see -[Startup and outage modes](#ingestion-modes). Create a deduplicated table -**before** ingestion if duplicates are unacceptable: a missing table is -auto-created without DEDUP. Keep event IDs and timestamps stable across -retries. +[Startup and outage modes](#ingestion-modes). Keep event IDs and timestamps +stable across retries. + +If duplicates are unacceptable, create a deduplicated table **before** the +first ingester runs, for example as a deployment or migration step. An +ingester that starts while QuestDB is down cannot create the table itself +first: when QuestDB comes back, the sender replays its journal in the +background, QuestDB auto-creates a missing table without DEDUP, and a later +`CREATE TABLE IF NOT EXISTS` does nothing. To add DEDUP to an existing table, +use [`ALTER TABLE ... DEDUP ENABLE`](/docs/query/sql/alter-table-enable-deduplication/), +for example +`ALTER TABLE trades_sf DEDUP ENABLE UPSERT KEYS(timestamp, trade_id);`. It +does not remove duplicates that were written earlier. @@ -606,10 +711,14 @@ queue. #### Lock recovery {#sf-lock-recovery} Node.js uses a `.lock.owner` directory in each journal slot, not an OS file -lock. A crashed process can leave one behind. If opening the journal fails -with `QwpReplayStoreLockedError` (wrapped in `QwpPoolResourceError` when -pooled), verify that no other process owns the slot **before** removing a -stale lock. See the +lock, and a crashed process can leave one behind. On the same host, the next +sender reclaims it automatically once the recorded process has exited, so a +restarted process recovers its slots without intervention. The client cannot +reclaim a lock recorded on another host, such as a container replaced under a +new host name, or one whose process ID now belongs to another running +process: opening the journal then fails with `QwpReplayStoreLockedError` +(wrapped in `QwpPoolResourceError` when pooled). Verify that no other process +owns the slot **before** removing such a lock; see the [Node.js lock-recovery runbook](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery). Do not let Node.js and another client's OS-lock-based sender use the same `sf_dir` concurrently. @@ -618,7 +727,13 @@ Do not let Node.js and another client's OS-lock-based sender use the same On QuestDB Enterprise with replication, `request_durable_ack=on` makes the acknowledgement watermark wait until the WAL has been uploaded to object -storage. `sender: { awaitDurableAck: true }` also makes each `flush()` wait: +storage. Typed `sender: { awaitDurableAck: true }` also makes each `flush()` +wait for the upload, for up to `durableAckTimeoutMs` (by default +`ackTimeoutMs`, 15 seconds). Under light load, the primary uploads WAL data +only when +[`replication.primary.throttle.window.duration`](/docs/high-availability/tuning/#throttle-window) +expires: 10 seconds by default, and 60 seconds in the network-efficiency +profile. Set the deadline well above it: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -627,7 +742,7 @@ const token = process.env.QDB_TOKEN; if (!token) throw new Error("QDB_TOKEN is not set"); const db = await connectQwpNodeClient( `wss::addr=db.example.com:9000;token=${token};request_durable_ack=on;`, - { sender: { awaitDurableAck: true } }, + { sender: { awaitDurableAck: true, durableAckTimeoutMs: 120_000 } }, ); try { const sender = await db.borrowSender(); @@ -637,7 +752,7 @@ try { .symbol("symbol", "ETH-USD") .doubleColumn("price", 2615.54) .at(Date.now(), "ms"); - await sender.flush(); // resolves once the batch is in object storage + await sender.flush(); // waits up to 2 minutes for the upload } finally { await sender.close(); } @@ -646,13 +761,23 @@ try { } ``` -If the server lacks support, a foreground start fails with -`QwpDurableAckUnavailableError` (wrapped in `QwpPoolResourceError` when -pooled). A sender with a background start retries from startup and emits -`durable-ack-unavailable` connection events **even with `sf_dir`**. With -`sf_dir` and a foreground start, only later mismatches, after a successful -connection, are retried. Monitor these events and journal capacity: a -successful background start does not prove durable ACK is available. +When the deadline passes, `flush()` rejects with a plain `Error` whose message +starts with `timed out waiting for QWP durable ACK`. QuestDB has already +acknowledged the batch, so it is committed to the WAL on the primary: the +timeout means only that the upload was not confirmed in time. + +If the server does not support durable ACK: + +- A sender with a foreground start fails with + `QwpDurableAckUnavailableError` (wrapped in `QwpPoolResourceError` when + pooled). +- A sender with a background start retries from startup and emits + `durable-ack-unavailable` connection events, **even with `sf_dir`**. +- A sender with `sf_dir` and a foreground start fails at its first + connection, but after a successful connection it retries later mismatches. + +Monitor these events and journal capacity: a successful background start +does not prove durable ACK is available. ### Fire-and-forget UDP @@ -986,12 +1111,15 @@ for large results; `compression_level=3` is accepted only with `zstd` or ## Error handling -Handle the failure at the stage where it occurs: +Handle the failure at the stage where it occurs. Every error class and +constant named on this page is exported from `@questdb/nodejs-client`, so you +can import it for `instanceof` checks: | Failure | Action | |---|---| | Local value validation | Fix the value; the row in progress was discarded. Test `QwpBatchTooLargeError` before `RangeError` because it extends `RangeError`. | | `QwpMemoryReplayAppendTimeoutError` / `QwpReplayStoreAppendTimeoutError` | The batch stays staged. Slow down and retry the flush, not the rows. | +| ACK timeout: `QwpIngressAckTimeoutError` from `waitForAcknowledged()`, or a plain `Error` from a `flush()` that waits for an acknowledgement | Not a rejection: QuestDB may still acknowledge the batch. Do not write the rows again; wait again, or raise `ackTimeoutMs` or `durableAckTimeoutMs`. See [Awaiting acknowledgements](#awaiting-acknowledgements). | | Server rejection | See [Ingestion errors](#ingestion-errors); a terminal rejection fails the sender. | | `QwpEgressQueryError` | Check `status`: fix SQL or bind values for `PARSE_ERROR`; retry on a new lease for `CANCELLED`, such as a query cancelled by a server shutdown. The lease remains usable after a SQL error. | | `QwpEgressSessionClosedError` | The query connection was lost with failover off. Close the lease and retry the whole query. | @@ -1018,6 +1146,12 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { toFsn: error.toFsn, }); }, + onError: (event) => { + // terminal: true means the sender has stopped; other events are warnings. + if (event.terminal) { + console.error("QuestDB ingestion stopped:", event.error.message); + } + }, }, }); await db.close(); @@ -1025,17 +1159,22 @@ await db.close(); Each error has `category`, `appliedPolicy`, `serverStatusByte`, `serverMessage`, `messageSequence`, the rejected frame range `fromFsn` to -`toFsn`, `tableName` when the server reports one, and `detectedAtMs`. +`toFsn`, `tableName` when the server reports one, `detectedAtMs`, and, when +abandoned store-and-forward data was preserved on disk, `quarantinedPath`. Categories and policies are lowercase, hyphenated strings (`QWP_SENDER_ERROR_CATEGORY`, `QWP_SENDER_ERROR_POLICY`). The handler also runs for retriable rejections, which the client resends: only `terminal` (the -sender stops) and `abandoned` (journal data set aside) mean the rows are not +sender stops) and `abandoned` (journal data set aside; see +[Quarantined journal slots](#quarantined-journal-slots)) mean the rows are not being delivered. The default policy of each category is listed under [Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). Without a callback, the client logs rejections. `waitForAcknowledged()` rejects for a rejected batch. Branch on category, not the unstable message -text, and redact messages before sending them to external trackers. Terminal -session errors also reach `ingressSession.onError` with `terminal: true`. +text, and redact messages before sending them to external trackers. +`ingressSession.onError` receives an event with `error`, `terminal`, +`timestampMs`, and, for a server rejection, `senderError`. It also reports +non-terminal problems, such as ACK timeouts; `terminal: true` means the sender +has stopped. #### Recovering from a terminal rejection @@ -1061,6 +1200,33 @@ journal: server-side fix, it replays the whole journal, including the rejected batch. After a move, it starts with an empty slot. +#### Quarantined journal slots + +Store-and-forward does not delete rows that it cannot deliver. It sets them +aside and reports them to `onSenderError` with category `data-loss` and +policy `abandoned`: + +- **Corrupt journal.** When a sender opens a slot whose journal is + structurally corrupt, the client renames the slot directory to + `.unreplayable-N`, adds a `.failed` file, and continues with an empty + slot. The error's `quarantinedPath` names the renamed directory. The client + never replays it; keep it for inspection. +- **Undeliverable slot.** A background drainer replays journal slots that no + running sender holds, including this client's own `-` slots. + When it cannot deliver a slot, for example because QuestDB terminally + rejects the oldest batch or rejects authentication, it adds a `.failed` + file to the slot and stops retrying it. The rows stay in the slot. + +To replay an undeliverable slot, fix the cause, then remove its marker with +`retryQwpNodeOrphanSlot()`. A running client replays the slot on its next +scan, within 30 seconds: + +```typescript +import { retryQwpNodeOrphanSlot } from "@questdb/nodejs-client"; + +await retryQwpNodeOrphanSlot("/var/lib/myapp/qdb-sf/trades-1"); +``` + ### Query errors SQL errors reject query iteration and `completion` with @@ -1192,12 +1358,10 @@ A query that fails over restarts from its first row; see Senders resend unacknowledged batches after a disconnect. A sender in default memory mode gives up after `reconnect_max_duration_millis` (5 minutes by default): it fails with `QwpReconnectExhaustedError` and its unacknowledged -rows are lost. Background memory mode or `sf_dir` retries indefinitely, -subject to queue or journal capacity. For the first connection, see -[Startup and outage modes](#ingestion-modes). Setting a `reconnect_*` key -implicitly requests bounded first-connection retry for senders unless you -explicitly set `initial_connect_retry=off`; `connectQwpNodeClient()` waits -for that retry only with `query_pool_min=0`. +rows are lost. Background memory mode and store-and-forward retry +indefinitely, subject to queue or journal capacity. Startup behavior and the +`reconnect_*` keys are covered in +[Startup and outage modes](#ingestion-modes). ### Query failover @@ -1247,15 +1411,20 @@ outages. When failover gives up, the query fails with ### Typed reconnect policy -Typed `ingressSession.reconnect` and `egressSession.reconnect` **replace** -their respective connect-string policies; set every limit you rely on in -the typed object. Ingestion `reconnect_*` keys trigger first-connection -retry, but a typed ingestion `reconnect` object does not. A query's first -connection retries only if `failover=on` is explicit, a `failover_*` key is -set without `failover=off`, or a typed query `reconnect` object is provided. -A typed `egressSession.reconnect` object also turns query failover back on -when the connect string says `failover=off`, so do not add one, even just for -`onEvent`, to a client that must not re-execute SQL. +Typed `ingressSession.reconnect` and `egressSession.reconnect` objects +**replace** their connect-string policies, so set every limit you rely on in +the typed object. + +- **Ingestion.** `reconnect_*` keys trigger first-connection retry; a typed + ingestion `reconnect` object does not. +- **Queries.** The first query connection is retried only if one of these is + set: + - `failover=on`, explicitly; + - a `failover_*` key, without `failover=off`; + - a typed `egressSession.reconnect` object. +- **Failover off.** A typed `egressSession.reconnect` object turns query + failover back on even when the connect string says `failover=off`. Do not + add one, even just for `onEvent`, to a client that must not re-execute SQL. ### Connection events @@ -1289,7 +1458,8 @@ const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { await db.close(); ``` -No event marks a terminal failure: use `ingressSession.onError` with +No event marks a terminal failure: use +[`ingressSession.onError`](#ingestion-errors) with `terminal: true` for that. An `egressSession.reconnect` object added for `onEvent` re-enables query failover on a `failover=off` client; see [Typed reconnect policy](#typed-reconnect-policy). Connect-string @@ -1339,7 +1509,7 @@ the | Area | Node.js behavior | |---|---| -| Startup/outage | `lazy_connect=on` or `initial_connect_retry=async` starts senders in the background, but `connectQwpNodeClient()` starts while QuestDB is down only with `query_pool_min=0`, which `lazy_connect=on` sets. First-connection retry from `initial_connect_retry=on` or a `reconnect_*` key covers senders only. Default memory mode gives up after 5 minutes per outage; SF and background memory modes retry indefinitely. See [Startup and outage modes](#ingestion-modes). | +| Startup/outage | `connectQwpNodeClient()` also opens a query connection, so it starts while QuestDB is down only with `query_pool_min=0`, which `lazy_connect=on` sets. A sender in default memory mode gives up after `reconnect_max_duration_millis` per outage. See [Startup and outage modes](#ingestion-modes). | | Authentication | After a sender's first successful connection, `401`/`403` is retried indefinitely by senders with `sf_dir` or a background start; other senders and queries treat it as terminal. | | Query startup | First-connect retry requires explicit `failover=on`, a `failover_*` key (without `failover=off`), or typed `egressSession.reconnect`. | | `target`, `zone` | Also apply to ingestion when set in the connect string; use typed `egress.target` for queries only. | diff --git a/documentation/high-availability/client-failover/concepts.md b/documentation/high-availability/client-failover/concepts.md index d4a2fc0cc9..93b154c190 100644 --- a/documentation/high-availability/client-failover/concepts.md +++ b/documentation/high-availability/client-failover/concepts.md @@ -57,7 +57,7 @@ host. | `Unknown` | The host has not been tried in this round, or its classification was reset. | | `TransientReject` | The server returned `421` with `X-QuestDB-Role: PRIMARY_CATCHUP` — it is a primary that is still catching up after promotion. Expected to recover. | | `TransportError` | TCP/TLS handshake failed, an HTTP upgrade returned a transient error code, or an established connection broke mid-stream. | -| `TopologyReject` | The server returned `421` with any role other than `PRIMARY_CATCHUP` (`PRIMARY`, `REPLICA`, `STANDALONE`, or an unrecognised token), or — on egress — a successfully-upgraded host whose `SERVER_INFO` role does not satisfy the requested `target=` filter. The host will not become usable without a topology change. | +| `TopologyReject` | The server returned `421` with any role other than `PRIMARY_CATCHUP` (`PRIMARY`, `REPLICA`, `STANDALONE`, or an unrecognised token), or a successfully-upgraded host whose `SERVER_INFO` role does not satisfy the requested `target=` filter (egress only; the Node.js client also applies it to ingress). The host will not become usable without a topology change. | A lower state in the table above is preferred when the client picks the next host to try. diff --git a/documentation/high-availability/client-failover/configuration.md b/documentation/high-availability/client-failover/configuration.md index 638aeeeea0..51596d8cce 100644 --- a/documentation/high-availability/client-failover/configuration.md +++ b/documentation/high-availability/client-failover/configuration.md @@ -65,7 +65,7 @@ network), and retrying for five minutes only hides it. |---|---| | `off` (default; alias `false`) | First-connect failure is terminal. The producer's call to build the sender throws immediately. | | `on` (aliases `sync`, `true`) | First-connect failures are retried on the caller's thread. The constructor blocks until it connects or `reconnect_max_duration_millis` expires — this is the **only** place that key applies. Once the sender is running, reconnection is unbounded, except for the Node.js senders described in the `reconnect_max_duration_millis` row above. | -| `async` | The constructor returns immediately; the background I/O thread drives the reconnect loop. The producer experiences backpressure if it tries to publish before the connection comes up. Intended for unattended producers where the SF directory may already carry segments from a prior process and the server may come up later. | +| `async` | The constructor returns immediately; the background I/O thread drives the reconnect loop. The producer experiences backpressure if it tries to publish before the connection comes up. Intended for unattended producers where the SF directory may already carry segments from a prior process and the server may come up later. On Node.js, `connectQwpNodeClient()` also opens a query connection at startup, so it returns while the server is down only with `query_pool_min=0`, which `lazy_connect=on` sets; see [startup and outage modes](/docs/connect/clients/nodejs/#ingestion-modes). | ## Egress (query) diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index bbf673e3e8..980cc1569b 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -121,9 +121,12 @@ is skipped, not stolen. ### Node.js lock recovery {#nodejs-lock-recovery} -A Node.js sender that crashed, or ran in a container that was replaced, can -leave its `.lock.owner` directory behind, and a new sender on the slot then -fails with `QwpReplayStoreLockedError`. To recover: +A Node.js sender can leave its `.lock.owner` directory behind when it crashes. +On the same host, the next sender reclaims the lock automatically once the +recorded process has exited. It cannot reclaim a lock recorded on another +host, such as a container replaced under a new host name, or one whose process +ID now belongs to another running process; a new sender on the slot then fails +with `QwpReplayStoreLockedError`. To recover: 1. Verify that the previous owner has exited and that no process is using the slot. The `.lock.owner` directory records the owner's host name and process From aa3e03009015414bee026908b52f88f728745552 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 18:28:28 +0100 Subject: [PATCH 23/25] docs(nodejs): fix lock recovery, drainer, and error guidance from review - Lock recovery: state the actual reclaim rule (same host, recorded process ID no longer in use) and that a container restarted in place keeps failing until the stale lock is removed. Add Node.js cautions to orphan adoption on the store-and-forward concepts and when-to-use pages. - Connection errors: with several endpoints, an authentication rejection surfaces as the QwpUpgradeError itself, not a QwpFailoverError. - Background drainer: describe its real scope (the client's own sender_id slots always, other sender_ids only with drain_orphans=on) in one place and link it from quarantine, connection events, the differences table, and the connect string reference. - Add a full example combining TLS, a token, two endpoints, store-and-forward, DEDUP, error callbacks, an ACK wait, and a failover-safe query. - List every sender error category with its default policy, and branch on the exported constants in the example. - Map typed reconnect fields to their connect-string keys and defaults, and list the reconnect and failover backoff keys inline. - Document writing from request handlers, the exported client classes, close() behavior per mode, staged rows dropped after an append timeout, and replays that return no batches. - QWP client behavior: include terminal server rejections in the async stop conditions, and correct the Java reconnect-loop appendix. --- .../connect/clients/connect-string.md | 5 + documentation/connect/clients/nodejs.md | 382 +++++++++++++++--- .../wire-protocols/qwp-client-behavior.md | 20 +- .../store-and-forward/concepts.md | 13 + .../store-and-forward/operating-and-tuning.md | 20 +- .../store-and-forward/when-to-use.md | 5 + 6 files changed, 374 insertions(+), 71 deletions(-) diff --git a/documentation/connect/clients/connect-string.md b/documentation/connect/clients/connect-string.md index 3c7a7adbc1..b015b14dbe 100644 --- a/documentation/connect/clients/connect-string.md +++ b/documentation/connect/clients/connect-string.md @@ -656,6 +656,11 @@ and releases it — **multiple orphans drain in parallel**, up to - `drain_orphans` — `on` enables the orphan drainer pool. Default: `off`. - `max_background_drainers` — maximum concurrent drainers. Default: `4`. +Without `drain_orphans=on`, the pooled Node.js client still replays slots of +its own `sender_id` that no running sender holds; the key adds other +`sender_id`s. See +[Node.js journal replay](/docs/connect/clients/nodejs/#replaying-the-journal-after-a-restart). + For delivery semantics, architecture, and tradeoffs (at-least-once guarantees, DEDUP requirements, segment-granular trim), see [Store-and-forward concepts](/docs/high-availability/store-and-forward/concepts/). diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index d985e3e275..8e38fd4a68 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -125,6 +125,10 @@ shared keys and [Differences from other clients](#differences-from-other-clients for Node.js exceptions. A second argument takes typed options; see [Programmatic options](#programmatic-options). +`connectQwpNodeClient()` resolves to a `QwpClient`. Its `borrowSender()` +returns a `QwpSender`, and its `borrowQuery()` a `QwpQueryLease`. The package +exports all three classes, so you can use them to type your own variables. + ### Standalone Sender For ingestion without a pool, create a `Sender` from a connect string, call @@ -258,13 +262,21 @@ How the keys combine: ### Closing the pooled client `db.close()` rejects new borrows, closes idle senders and queries, and waits -briefly for borrowed senders to be returned. It can resolve without every -batch being acknowledged. Wait for `sender.publishedSequence` before returning -a borrowed sender when an ACK is required, or use `sf_dir` to retain unacked -rows across restarts. In [default memory mode](#ingestion-modes), a borrowed -sender's `close()` can wait for a reconnect up to -`reconnect_max_duration_millis` (5 minutes by -default); plan your shutdown deadline accordingly. +briefly for borrowed senders to be returned. It then waits up to +`close_flush_timeout_millis` (5 seconds by default) for acknowledgements and +resolves even if some batches are still unacknowledged: with `sf_dir`, they +stay in the journal and are replayed on the next start; without it, they are +lost. When an ACK is required, call +`await sender.waitForAcknowledged(sender.publishedSequence, timeoutMs)` before +returning a borrowed sender. + +A borrowed sender's `close()` flushes its completed rows. With a background +start or `sf_dir`, that hands them to the memory replay queue or the journal, +so `close()` returns without waiting for QuestDB, even during an outage, +unless the queue or journal is full (see [Backpressure](#backpressure)). In +[default memory mode](#ingestion-modes), `close()` can wait for a reconnect up +to `reconnect_max_duration_millis` (5 minutes by default); plan your shutdown +deadline accordingly. ## Data ingestion @@ -457,8 +469,10 @@ The replay queue defaults to 128 MiB without `sf_dir`, and the disk journal targets 10 GiB with it. When full, publishing waits up to 30 seconds by default, then rejects with `QwpMemoryReplayAppendTimeoutError` or `QwpReplayStoreAppendTimeoutError`. The batch stays staged: slow down and -retry `flush()`; do not write the rows again. Configure the cap with -`sf_max_total_bytes` and the wait with `sf_append_deadline_millis`. +retry `flush()`; do not write the rows again. If you close the sender +instead, `close()` tries once more and, if there is still no room, drops the +staged rows with a warning and rejects with the same error. Configure the cap +with `sf_max_total_bytes` and the wait with `sf_append_deadline_millis`. #### Batch size limits @@ -660,9 +674,13 @@ and [operating guide](/docs/high-availability/store-and-forward/operating-and-tu #### Replaying the journal after a restart Run a client with the same `sf_dir` and `sender_id` once QuestDB is reachable -again. The pool's first sender reopens the journal slot (`trades-0` here) at -startup and replays the unacknowledged frames in the background, so keep the -default `sender_pool_min=1`. Then poll for a row you know was written: +again. It replays the unacknowledged frames in the background, whatever +`sender_pool_min` is: each pooled sender reopens its own slot (`trades-0` for +the first), and a background drainer replays every `trades-` slot that no +running sender holds, at startup and then every 30 seconds. That includes +slots left by more concurrent senders in an earlier run. Slots of other +`sender_id`s in the same `sf_dir` are replayed only with `drain_orphans=on`. +Then poll for a row you know was written: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; @@ -711,15 +729,19 @@ queue. #### Lock recovery {#sf-lock-recovery} Node.js uses a `.lock.owner` directory in each journal slot, not an OS file -lock, and a crashed process can leave one behind. On the same host, the next -sender reclaims it automatically once the recorded process has exited, so a -restarted process recovers its slots without intervention. The client cannot -reclaim a lock recorded on another host, such as a container replaced under a -new host name, or one whose process ID now belongs to another running -process: opening the journal then fails with `QwpReplayStoreLockedError` -(wrapped in `QwpPoolResourceError` when pooled). Verify that no other process -owns the slot **before** removing such a lock; see the -[Node.js lock-recovery runbook](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery). +lock, and a crashed process can leave one behind. The next sender reclaims it +automatically only when the lock was recorded on the same host and the +recorded process ID is no longer in use, as when a process restarts on the +same machine under a new process ID. It cannot reclaim a lock recorded on +another host, such as a container replaced under a new host name, or one whose +process ID is in use again. That includes a container restarted in place, +which usually gives the restarted process its previous process ID (often 1), +so the restarted process holds the recorded ID itself. Opening the journal +then fails with `QwpReplayStoreLockedError` (wrapped in `QwpPoolResourceError` +when pooled) on every start until the stale lock is removed. Verify that no +other process owns the slot **before** removing it; the +[Node.js lock-recovery runbook](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery) +also shows when a startup step can remove it safely. Do not let Node.js and another client's OS-lock-based sender use the same `sf_dir` concurrently. @@ -1129,22 +1151,36 @@ can import it for `instanceof` checks: A local column or `at()` validation error discards the unfinished row. A **server** rejection can arrive after `flush()` resolves: register -`ingressSession.onSenderError` to receive it. +`ingressSession.onSenderError` to receive it, and branch on the error's +`appliedPolicy` and `category`. ```typescript -import { connectQwpNodeClient } from "@questdb/nodejs-client"; +import { + connectQwpNodeClient, + QWP_SENDER_ERROR_CATEGORY, + QWP_SENDER_ERROR_POLICY, +} from "@questdb/nodejs-client"; const db = await connectQwpNodeClient("ws::addr=localhost:9000;", { ingressSession: { onSenderError: (error) => { // Keep serverMessage out of external trackers: it may contain row values. - console.error("QuestDB rejected a batch", { - category: error.category, // for example "schema-mismatch" - policy: error.appliedPolicy, // "retriable", "retriable-other", "terminal", or "abandoned" + const details = { + category: error.category, table: error.tableName, fromFsn: error.fromFsn, toFsn: error.toFsn, - }); + }; + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + // The sender has stopped; see "Recovering from a terminal rejection". + console.error("QuestDB rejected a batch; ingestion stopped", details); + } else if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.ABANDONED) { + console.error("rows set aside in", error.quarantinedPath, details); + } else if (error.category === QWP_SENDER_ERROR_CATEGORY.WRITE_ERROR) { + console.warn("QuestDB could not write a batch; resending", details); + } else { + console.warn("QuestDB rejected a batch; resending", details); + } }, onError: (event) => { // terminal: true means the sender has stopped; other events are warnings. @@ -1161,14 +1197,29 @@ Each error has `category`, `appliedPolicy`, `serverStatusByte`, `serverMessage`, `messageSequence`, the rejected frame range `fromFsn` to `toFsn`, `tableName` when the server reports one, `detectedAtMs`, and, when abandoned store-and-forward data was preserved on disk, `quarantinedPath`. -Categories and policies are lowercase, hyphenated strings -(`QWP_SENDER_ERROR_CATEGORY`, `QWP_SENDER_ERROR_POLICY`). The handler also -runs for retriable rejections, which the client resends: only `terminal` (the -sender stops) and `abandoned` (journal data set aside; see -[Quarantined journal slots](#quarantined-journal-slots)) mean the rows are not -being delivered. The default policy of each category is listed under -[Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames). -Without a callback, the client logs rejections. `waitForAcknowledged()` +Categories and policies are lowercase, hyphenated strings; compare them with +the `QWP_SENDER_ERROR_CATEGORY` and `QWP_SENDER_ERROR_POLICY` constants. Each +category has a fixed default policy, because Node.js does not apply the +`on_*_error` keys: + +| Category | Default policy | Meaning | +|---|---|---| +| `schema-mismatch` | `terminal` | The batch does not match the table schema | +| `parse-error` | `terminal` | QuestDB could not parse the batch | +| `security-error` | `terminal` | QuestDB denied the write, for example by ACL | +| `protocol-violation` | `terminal` | The client and server disagree on the protocol | +| `write-error` | `retriable` | The write failed, for example on a table that is not accepting writes | +| `internal-error` | `retriable` | An unexpected server-side failure | +| `dictionary-gap` | `retriable` | The connection lacks symbol dictionary entries; the client resends them | +| `not-writable` | `retriable-other` | The node cannot accept writes; the client tries another endpoint | +| `cancelled`, `limit-exceeded` | `retriable` | Current servers send these statuses only on query connections | +| `unknown` | `retriable` | A status this client does not know, for example from a newer server | +| `data-loss` | `abandoned` | Store-and-forward data was set aside; see [Quarantined journal slots](#quarantined-journal-slots) | + +The handler also runs for retriable rejections, which the client resends: only +`terminal` (the sender stops) and `abandoned` (journal data set aside) mean +the rows are not being delivered. Without a callback, the client logs +rejections. `waitForAcknowledged()` rejects for a rejected batch. Branch on category, not the unstable message text, and redact messages before sending them to external trackers. `ingressSession.onError` receives an event with `error`, `terminal`, @@ -1211,11 +1262,11 @@ policy `abandoned`: `.unreplayable-N`, adds a `.failed` file, and continues with an empty slot. The error's `quarantinedPath` names the renamed directory. The client never replays it; keep it for inspection. -- **Undeliverable slot.** A background drainer replays journal slots that no - running sender holds, including this client's own `-` slots. - When it cannot deliver a slot, for example because QuestDB terminally - rejects the oldest batch or rejects authentication, it adds a `.failed` - file to the slot and stops retrying it. The rows stay in the slot. +- **Undeliverable slot.** When the background drainer (see + [Replaying the journal after a restart](#replaying-the-journal-after-a-restart)) + cannot deliver a slot, for example because QuestDB terminally rejects the + oldest batch or rejects authentication, it adds a `.failed` file to the + slot and stops retrying it. The rows stay in the slot. To replay an undeliverable slot, fix the cause, then remove its marker with `retryQwpNodeOrphanSlot()`. A running client replays the slot on its next @@ -1279,8 +1330,11 @@ WebSocket, so check its `kind`: `authentication` for an HTTP `401` or `403`, `transport` or `timeout` when QuestDB is unreachable, and others such as `role-rejected` and `version-mismatch`. `QwpRoleMismatchError` and `QwpDurableAckUnavailableError` extend `QwpUpgradeError`, so test for them -first. With several endpoints, the cause is a `QwpFailoverError` whose -`attempts` records each endpoint's failure; with one endpoint, it is that +first. With several endpoints, a connection that fails on every endpoint, for +example because none is reachable, has a `QwpFailoverError` cause whose +`attempts` records each endpoint's failure. An authentication rejection is the +exception: it stops the endpoint sweep at once, so the cause is that +`QwpUpgradeError` itself. With one endpoint, the cause is always that endpoint's error. A borrow at pool capacity times out as `QwpPoolAcquireTimeoutError` after `acquire_timeout_ms` (5 seconds by default). @@ -1355,19 +1409,23 @@ A query that fails over restarts from its first row; see ### Ingestion reconnect -Senders resend unacknowledged batches after a disconnect. A sender in default -memory mode gives up after `reconnect_max_duration_millis` (5 minutes by -default): it fails with `QwpReconnectExhaustedError` and its unacknowledged -rows are lost. Background memory mode and store-and-forward retry -indefinitely, subject to queue or journal capacity. Startup behavior and the -`reconnect_*` keys are covered in +Senders resend unacknowledged batches after a disconnect. Between attempts, +they wait a random delay below a ceiling that starts at +`reconnect_initial_backoff_millis` (100 ms by default) and doubles up to +`reconnect_max_backoff_millis` (5 seconds). A sender in default memory mode +gives up after `reconnect_max_duration_millis` (5 minutes by default): it +fails with `QwpReconnectExhaustedError` and its unacknowledged rows are lost. +Background memory mode and store-and-forward retry indefinitely, subject to +queue or journal capacity. Setting any `reconnect_*` key also makes senders +retry their first connection; see [Startup and outage modes](#ingestion-modes). ### Query failover Query failover is on by default. A lost query connection re-executes the -query **from its first row**, even if your loop has already consumed rows, and -a replay can also return zero batches. For streaming results you cannot +query **from its first row**, even if your loop has already consumed rows. The +re-executed query reads the data as it is then, so it can return fewer rows +than the first attempt, or no batches at all. For streaming results you cannot retract, use `failover=off` and retry the whole operation when the query fails with `QwpEgressSessionClosedError`. Otherwise buffer the result and reset the buffer in `egressSession.onReplayReset`. Give that callback a client with one @@ -1404,16 +1462,36 @@ try { ``` `batch.batchSequence === 0n` detects a nonempty replay but not a replay -returning no batches. Query failover defaults to 8 attempts, which may end -before the 30-second time budget: raise `failover_max_attempts` for longer -outages. When failover gives up, the query fails with -`QwpReconnectExhaustedError`; close the lease and borrow a new one. +returning no batches. A single-row aggregate, such as `count()` or `avg()` +without `GROUP BY`, needs no reset: every execution returns exactly one row, +so keep the last row you receive. + +Between failover attempts, the client waits a random delay below a ceiling +that starts at `failover_backoff_initial_ms` (50 ms by default) and doubles up +to `failover_backoff_max_ms` (1 second). It makes at most +`failover_max_attempts` (8) attempts within `failover_max_duration_ms` +(30 seconds), so the attempt limit can end failover before the time budget: +raise `failover_max_attempts` for longer outages. When failover gives up, the +query fails with `QwpReconnectExhaustedError`; close the lease and borrow a +new one. ### Typed reconnect policy Typed `ingressSession.reconnect` and `egressSession.reconnect` objects -**replace** their connect-string policies, so set every limit you rely on in -the typed object. +**replace** their connect-string policies: a field you omit takes the default +below, not the connect-string value, so set every limit you rely on in the +typed object. Durations are in milliseconds, and `maxDurationMs: 0` removes +the time limit. + +| Field | Ingestion key (default) | Query key (default) | +|---|---|---| +| `initialBackoffMs` | `reconnect_initial_backoff_millis` (100) | `failover_backoff_initial_ms` (50) | +| `maxBackoffMs` | `reconnect_max_backoff_millis` (5000) | `failover_backoff_max_ms` (1000) | +| `maxDurationMs` | `reconnect_max_duration_millis` (300000) | `failover_max_duration_ms` (30000) | +| `maxAttempts` | No key (0, unlimited) | `failover_max_attempts` (8) | +| `maxFrameRejections` | `max_frame_rejections` (4) | Not used | +| `poisonMinEscalationWindowMs` | `poison_min_escalation_window_millis` (300000) | Not used | +| `onEvent` | No key; see [Connection events](#connection-events) | No key | - **Ingestion.** `reconnect_*` keys trigger first-connection retry; a typed ingestion `reconnect` object does not. @@ -1433,8 +1511,10 @@ observe connection events. Each event has a `kind`, an `attempt` number, `timestampMs`, and, where relevant, `endpoint`, `previousEndpoint`, and `cause`. The kinds are `connected`, `reconnecting`, `attempt-failed` (one per failed connection attempt, with its `cause`), `reconnected`, `failed-over`, -and `durable-ack-unavailable`. Orphan drainers (`drain_orphans=on`) also -report `primary-unavailable` and `durable-ack-persistent-failure`. +and `durable-ack-unavailable`. The store-and-forward background drainer (see +[Replaying the journal after a restart](#replaying-the-journal-after-a-restart)) +sends its events to the same `ingressSession.reconnect.onEvent` and also +reports `primary-unavailable` and `durable-ack-persistent-failure`. `QWP_RECONNECT_EVENT_KIND` lists them all. ```typescript @@ -1472,6 +1552,77 @@ Share one `QwpClient`, but keep one sender per producer and one query lease per concurrent query. Worker threads need their own clients and, with `sf_dir`, distinct `sender_id` values. +### Writing from request handlers + +In a Node.js service, every request handler that writes rows is a concurrent +producer, even though all handlers run on one thread. Two patterns work: + +- **Borrow per request.** `borrowSender()` hands out an idle pooled sender + without reconnecting, and `close()` flushes the request's rows and returns + the sender. At most `sender_pool_max` handlers (4 by default) hold a sender + at once; another borrow waits up to `acquire_timeout_ms` (5 seconds by + default), then fails with `QwpPoolAcquireTimeoutError`. Each request sends + its own batch. With `sf_dir`, each pooled sender journals into its own + `-` slot. +- **One shared sender.** Borrow one sender at startup and build each row in + one synchronous chain from `table()` to `at()` or `atNow()`, with no `await` + in between. Rows from concurrent handlers then never interleave, auto-flush + batches them together, and flushes are serialized. A handler that awaits + mid-row makes the next handler's `table()` throw. Auto-flush runs only when a + row is added, so also call `flush()` from a timer, or the last rows wait for + the next request. A terminal failure stops the sender for every handler, + and with `transaction=on` all handlers share one transaction. + +With a background start, `borrowSender()`, `at()`, `flush()`, and `close()` +return promptly while QuestDB is down, until the replay queue or journal is +full (see [Backpressure](#backpressure)); only `borrowQuery()` fails until +QuestDB is reachable. In default memory mode, `flush()`, `close()`, and an +auto-flushing `at()` wait for the reconnect instead; see +[Startup and outage modes](#ingestion-modes). + +```typescript +import { createServer } from "node:http"; +import { connectQwpNodeClient } from "@questdb/nodejs-client"; + +// A background start lets the service start, and record trades, while +// QuestDB is down. +const db = await connectQwpNodeClient( + "ws::addr=localhost:9000;lazy_connect=on;sf_max_segment_bytes=1m;", +); + +async function recordTrade(price: number, amount: number): Promise { + const sender = await db.borrowSender(); + try { + await sender + .table("trades") + .symbol("symbol", "ETH-USD") + .symbol("side", "buy") + .doubleColumn("price", price) + .doubleColumn("amount", amount) + .at(Date.now(), "ms"); + } finally { + await sender.close(); // flushes the row and returns the sender + } +} + +const server = createServer((req, res) => { + recordTrade(2615.54, 0.5).then( + () => res.end("recorded\n"), + (error) => { + console.error("could not record the trade:", error); + res.statusCode = 503; + res.end(); + }, + ); +}); +server.listen(8080); + +process.once("SIGTERM", () => { + // db.close() waits up to close_flush_timeout_millis for acknowledgements. + server.close(() => void db.close()); +}); +``` + ## Configuration reference @@ -1514,6 +1665,7 @@ the | Query startup | First-connect retry requires explicit `failover=on`, a `failover_*` key (without `failover=off`), or typed `egressSession.reconnect`. | | `target`, `zone` | Also apply to ingestion when set in the connect string; use typed `egress.target` for queries only. | | `sf_dir` | Node.js recursively creates missing parents and a slot; pooled senders use slots named `-`. Its `.lock.owner` directory can outlive a crashed process; see [Lock recovery](#sf-lock-recovery). | +| Background drainer | With `sf_dir`, a pooled client replays slots of its own `sender_id` that no running sender holds, even without `drain_orphans`; `drain_orphans=on` adds other `sender_id`s. See [Replaying the journal after a restart](#replaying-the-journal-after-a-restart). | | SF-only keys | Explicit `sf_durability` (even `memory`), `sf_sync_interval_millis`, `drain_orphans`, `max_background_drainers`, and `catch_up_cap_gap_min_escalation_window_millis` require `sf_dir`. `sf_durability=append` is supported. | | `sf_max_total_bytes` | With `sf_dir` it is a journal size **target**, not a disk quota; without `sf_dir` it caps the memory queue. | | Durable ACK | Background-started senders, including with `sf_dir`, retry an unavailable capability from startup. Explicit `durable_ack_keepalive_interval_millis` also requests durable ACK even at `0`; negatives are rejected. | @@ -1550,6 +1702,120 @@ values); `decimalColumn()` rejects non-integer scales instead of silently coercing them. The package adds `ws` for QWP and retains the old ILP transports. +## Full example: ingestion and querying with failover + +This program combines the production settings from the sections above: TLS +and a token, two endpoints, a background start with store-and-forward, +deduplicated replays, error callbacks, an acknowledgement wait, and a query +that is safe under failover. If QuestDB is unreachable, the program still +starts and journals the rows, and `borrowQuery()` fails once query failover +gives up. Create the table first, for example as a migration step, so that +replays are deduplicated even if an ingester starts while QuestDB is down: + +```questdb-sql +CREATE TABLE IF NOT EXISTS trades_sf ( + timestamp TIMESTAMP, + trade_id VARCHAR, + symbol SYMBOL, + price DOUBLE +) TIMESTAMP(timestamp) PARTITION BY DAY +DEDUP UPSERT KEYS(timestamp, trade_id); +``` + +```typescript +import { + connectQwpNodeClient, + QWP_SENDER_ERROR_POLICY, + QwpIngressAckTimeoutError, +} from "@questdb/nodejs-client"; + +const token = process.env.QDB_TOKEN; +if (!token) throw new Error("QDB_TOKEN is not set"); + +const db = await connectQwpNodeClient( + "wss::addr=db-a.example.com:9000,db-b.example.com:9000;" + + `token=${token};` + + // Start while QuestDB is down and journal rows across restarts. + "lazy_connect=on;sf_dir=/var/lib/myapp/qdb-sf;sender_id=trades;" + + "sf_max_segment_bytes=1m;" + + // Bound query buffering; let query failover run for up to a minute. + "initial_credit=1048576;failover_max_duration_ms=60000;", + { + ingressSession: { + onSenderError: (error) => { + if (error.appliedPolicy === QWP_SENDER_ERROR_POLICY.TERMINAL) { + console.error("ingestion stopped:", error.category, error.tableName); + } + }, + onError: (event) => { + if (event.terminal) console.error("ingestion stopped:", event.error); + }, + }, + }, +); +try { + // Stable trade IDs and event timestamps let DEDUP absorb replays. + const now = Date.now(); + const fills = [ + { tradeId: "trade-1001", symbol: "ETH-USD", price: 2615.54, tsMs: now }, + { tradeId: "trade-1002", symbol: "ETH-USD", price: 2615.62, tsMs: now + 1 }, + ]; + const sender = await db.borrowSender(); + try { + for (const fill of fills) { + await sender + .table("trades_sf") + .stringColumn("trade_id", fill.tradeId) + .symbol("symbol", fill.symbol) + .doubleColumn("price", fill.price) + .at(fill.tsMs, "ms"); + } + await sender.flush(); // journaled locally + try { + await sender.waitForAcknowledged(sender.publishedSequence, 10_000); + } catch (error) { + if (!(error instanceof QwpIngressAckTimeoutError)) throw error; + // Not a rejection: the journal keeps the rows and replays them. + console.warn("not acknowledged yet; the rows stay in the journal"); + } + } finally { + await sender.close(); + } + + // A single-row aggregate is safe under query failover: a replay returns + // its own row, which replaces the first attempt's. + const lease = await db.borrowQuery(); + try { + const sinceMicros = BigInt(Date.now() - 3_600_000) * 1000n; + const query = await lease.query( + "SELECT count(), avg(price) FROM trades_sf " + + "WHERE symbol = $1 AND timestamp >= $2", + { + binds: (binds) => + binds.setVarchar(0, "ETH-USD").setTimestampMicros(1, sinceMicros), + timeoutMs: 30_000, + }, + ); + let summary: unknown[] = []; + for await (const batch of query) { + for (const row of batch.rows()) summary = [...row]; + } + await query.completion; + // Rows written above may not be visible yet; see Read-after-write. + console.log("ETH-USD, last hour [trades, average price]:", summary); + } finally { + await lease.close(); + } +} finally { + await db.close(); +} +``` + +Keep `target=replica` out of this connect string: on Node.js it also filters +ingestion; see [Multiple endpoints](#multiple-endpoints). To also wait for the +upload to object storage on QuestDB Enterprise, see +[Durable acknowledgement](#durable-acknowledgement). + ## ILP transports (legacy) The standalone `Sender` still speaks InfluxDB Line Protocol (ILP) over HTTP diff --git a/documentation/connect/wire-protocols/qwp-client-behavior.md b/documentation/connect/wire-protocols/qwp-client-behavior.md index c7e8c807ea..7ee1129844 100644 --- a/documentation/connect/wire-protocols/qwp-client-behavior.md +++ b/documentation/connect/wire-protocols/qwp-client-behavior.md @@ -361,9 +361,12 @@ With `initial_connect_retry=async`: surface on later producer calls or at close-time. A sender in async mode does not give up because time passed. What stops it is -a terminal condition: poison-frame escalation at any time, and an -authentication rejection or durable-ack capability mismatch before its first -successful connection. After that first connection, the Java reference client +a terminal condition: a server rejection with a terminal policy (by default +`SCHEMA_MISMATCH`, `PARSE_ERROR`, `SECURITY_ERROR`, and `PROTOCOL_VIOLATION`; +see [Error frames](/docs/high-availability/store-and-forward/concepts/#error-frames)), +poison-frame escalation at any time, and an authentication rejection or +durable-ack capability mismatch before its first successful connection. After +that first connection, the Java reference client retries authentication and durable-ack rejections instead, so a credential or capability change on the cluster cannot stop the producer. The Rust, C, C++, Python, Go, and .NET clients treat them as terminal. Producer calls can also @@ -595,5 +598,12 @@ source states the contract directly: > converts it into the durable-ack capability-gap budget. Neither bounds this > loop's steady-state reconnect. -`QwpAuthFailedException` and `WebSocketUpgradeException` raised inside the loop -are terminal across all endpoints. Everything else is retried. +`QwpAuthFailedException`, `WebSocketUpgradeException`, and +`QwpDurableAckMismatchException` raised inside the loop are terminal across +all endpoints before the sender's first successful connection, and in an +orphan drainer. After a first successful connection, the loop reports each one +to the error handler as `RETRIABLE` (`SECURITY_ERROR` for an authentication or +upgrade rejection, `PROTOCOL_VIOLATION` for a durable-ack mismatch) and keeps +retrying; see +[Authentication is cluster-wide](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide). +Everything else is retried. diff --git a/documentation/high-availability/store-and-forward/concepts.md b/documentation/high-availability/store-and-forward/concepts.md index de7e60633d..e98647f69e 100644 --- a/documentation/high-availability/store-and-forward/concepts.md +++ b/documentation/high-availability/store-and-forward/concepts.md @@ -338,6 +338,19 @@ until an operator intervenes. The orphan flow is opt-in because in a multi-tenant deployment with shared `sf_dir`, blindly draining unknown slots may be surprising. +:::caution Node.js client + +A Node.js drainer adopts a slot only if it can take the slot's `.lock.owner` +lock. It can reclaim a crashed owner's lock only on the same host, and only +when the recorded process ID is no longer in use, so it skips a slot whose +owner ran in a replaced container or in a container restarted in place. Those +rows stay on disk until the stale lock is removed; see +[Node.js lock recovery](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery). +A pooled Node.js client also replays slots of its own `sender_id` without +`drain_orphans=on`; the key adds the other `sender_id`s. + +::: + ## Error frames Not every server response is an OK. A rejected batch is **not** silently diff --git a/documentation/high-availability/store-and-forward/operating-and-tuning.md b/documentation/high-availability/store-and-forward/operating-and-tuning.md index 980cc1569b..d4a08b3b21 100644 --- a/documentation/high-availability/store-and-forward/operating-and-tuning.md +++ b/documentation/high-availability/store-and-forward/operating-and-tuning.md @@ -60,9 +60,11 @@ The Node.js client does not use an OS lock. It locks a slot with a `.lock.owner` directory that records the owner's host name and process ID, and keeps `.lock` and `.lock.pid` only for compatibility. After a crash, a new Node.js sender takes the slot over automatically only on the same host, once -the recorded process ID is no longer in use. In containers that usually fails, -because the application runs as process ID 1 and a replacement container has a -new host name, and the new sender reports `QwpReplayStoreLockedError`. +the recorded process ID is no longer in use. Containers usually defeat that +check: a replacement container has a new host name, and a container restarted +in place typically gives the restarted process its previous process ID, which +is often 1. The new sender then reports `QwpReplayStoreLockedError` on every +start. [Node.js lock recovery](#nodejs-lock-recovery) describes how to remove a stale lock safely. @@ -123,10 +125,11 @@ is skipped, not stolen. A Node.js sender can leave its `.lock.owner` directory behind when it crashes. On the same host, the next sender reclaims the lock automatically once the -recorded process has exited. It cannot reclaim a lock recorded on another -host, such as a container replaced under a new host name, or one whose process -ID now belongs to another running process; a new sender on the slot then fails -with `QwpReplayStoreLockedError`. To recover: +recorded process ID is no longer in use. It cannot reclaim a lock recorded on +another host, such as a container replaced under a new host name, or one whose +process ID is in use again, including by the restarted process itself in a +container restarted in place; a new sender on the slot then fails with +`QwpReplayStoreLockedError`. To recover: 1. Verify that the previous owner has exited and that no process is using the slot. The `.lock.owner` directory records the owner's host name and process @@ -143,7 +146,8 @@ Never delete the shared `.slot-locks` directory or another slot's locks. Automate this cleanup only where the deployment guarantees that the previous owner has exited before a new one starts, for example a single replica that -uses the Kubernetes `Recreate` update strategy and a `ReadWriteOnce` volume. A +uses the Kubernetes `Recreate` update strategy and a `ReadWriteOnce` volume, +or a container restarted in place whose volume no other process uses. A startup step can then remove the stale owner directories of the client's own slots before it creates the client. Anywhere two processes can overlap, recover manually. See also the diff --git a/documentation/high-availability/store-and-forward/when-to-use.md b/documentation/high-availability/store-and-forward/when-to-use.md index 6f4059043d..fa9a40ce34 100644 --- a/documentation/high-availability/store-and-forward/when-to-use.md +++ b/documentation/high-availability/store-and-forward/when-to-use.md @@ -154,6 +154,11 @@ spawn background drainers to clear them. - You prefer "automatic eventual delivery" over "operator manually reattaches the slot." +On the Node.js client, both a restarted process and a drainer recover a +crashed sender's slot only if they can reclaim its lock, which fails in a +replaced container or in a container restarted in place. Clear such locks +with [Node.js lock recovery](/docs/high-availability/store-and-forward/operating-and-tuning/#nodejs-lock-recovery). + ### Leave it off when - Each `sender_id` is statically pinned to a specific process — there From e1628246d50a802e6f4a62ac5ed1382059398894 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 22:02:27 +0100 Subject: [PATCH 24/25] docs(replication): correct the throttle window default to 1 second QuestDB Enterprise 3.3.1 lowered the default of replication.primary.throttle.window.duration from 10 seconds to 1 second. - Replication tuning: update the settings table, the default profile, the cost table's default row, the code sample, and the value table - Replication configuration reference: default 1000 - Node.js client: the durable acknowledgement section quoted the old default; link the 60-second figure to the network-efficiency settings - Changelog: note the correction --- documentation/changelog.mdx | 1 + .../configuration/database-replication.md | 2 +- documentation/connect/clients/nodejs.md | 5 +++-- documentation/high-availability/tuning.md | 14 +++++++------- 4 files changed, 12 insertions(+), 10 deletions(-) diff --git a/documentation/changelog.mdx b/documentation/changelog.mdx index 3f676f3054..f4b2fb6dc7 100644 --- a/documentation/changelog.mdx +++ b/documentation/changelog.mdx @@ -14,6 +14,7 @@ This page tracks significant updates to the QuestDB documentation. - [Node.js client](/docs/connect/clients/nodejs/) - Rewrote the page for QWP support in `@questdb/nodejs-client` 5.0.0: pooled ingestion and streaming SQL queries from one connect string, every column type, compiled object-row writers, acknowledgements, transactions, store-and-forward, UDP, failover, error handling, and migration from ILP and from 4.x. The [connect string reference](/docs/connect/clients/connect-string/) and the high-availability and wire-protocol pages now note where the Node.js client's keys, defaults, and behavior differ - [Store-and-forward](/docs/high-availability/store-and-forward/concepts/) - Corrected the replay semantics for every client: replay is at least once and can insert duplicate rows unless the table uses `DEDUP UPSERT KEYS`. The error policy table now shows the real defaults, which include no drop policy, and the [Java](/docs/connect/clients/java/#ingestion-errors) and [.NET](/docs/connect/clients/dotnet/#how-errors-surface) client pages now document their retriable, terminal, and abandoned policies. [Client failover](/docs/high-availability/client-failover/concepts/#authentication-is-cluster-wide) now lists which clients retry authentication rejections after a sender's first connection +- [Replication tuning](/docs/high-availability/tuning/) - Corrected the default of `replication.primary.throttle.window.duration` to 1 second, its value since QuestDB Enterprise 3.3.1, in the tuning guide and the [replication configuration reference](/docs/configuration/database-replication/#replicationprimarythrottlewindowduration) ## September 2026 diff --git a/documentation/configuration/database-replication.md b/documentation/configuration/database-replication.md index d50b9b9105..fef85cb1ab 100644 --- a/documentation/configuration/database-replication.md +++ b/documentation/configuration/database-replication.md @@ -136,7 +136,7 @@ better for constrained networks but more costly. ### replication.primary.throttle.window.duration -- **Default**: `10000` +- **Default**: `1000` - **Reloadable**: no The millisecond duration of the sliding window used to process replication diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 8e38fd4a68..2a952460ac 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -754,8 +754,9 @@ wait for the upload, for up to `durableAckTimeoutMs` (by default `ackTimeoutMs`, 15 seconds). Under light load, the primary uploads WAL data only when [`replication.primary.throttle.window.duration`](/docs/high-availability/tuning/#throttle-window) -expires: 10 seconds by default, and 60 seconds in the network-efficiency -profile. Set the deadline well above it: +expires: 1 second by default, or 60 seconds with the +[network-efficiency settings](/docs/high-availability/tuning/#network-efficiency). +Set the deadline well above the configured window: ```typescript import { connectQwpNodeClient } from "@questdb/nodejs-client"; diff --git a/documentation/high-availability/tuning.md b/documentation/high-availability/tuning.md index 73b93a1a9b..0b34977b2e 100644 --- a/documentation/high-availability/tuning.md +++ b/documentation/high-availability/tuning.md @@ -22,7 +22,7 @@ restart. | Setting | Node | Default | What it does | |---------|------|---------|-------------| -| `replication.primary.throttle.window.duration` | Primary | `10000` (10s) | Maximum time before an incomplete WAL segment is flushed | +| `replication.primary.throttle.window.duration` | Primary | `1000` (1s) | Maximum time before an incomplete WAL segment is flushed | | `replication.replica.poll.interval` | Replica | `1000` (1s) | How often the replica checks for new data | | `cairo.wal.segment.rollover.size` | Primary | `2097152` (2 MiB) | Max WAL segment size before rollover | @@ -59,7 +59,7 @@ replication.replica.poll.interval=100 No configuration needed. The defaults are: -- `replication.primary.throttle.window.duration=10000` (10s) +- `replication.primary.throttle.window.duration=1000` (1s) - `replication.replica.poll.interval=1000` (1s) - `cairo.wal.segment.rollover.size=2097152` (2 MiB) @@ -124,8 +124,8 @@ write ops typically cost ~\$5/million and read ops ~\$0.40/million. |---|---| | 50ms / 50ms | ~$280 | | 100ms / 100ms | ~$140 | -| 1s / 1s | ~$14 | -| 10s / 1s (default) | ~$2 | +| 1s / 1s (default) | ~$14 | +| 10s / 1s | ~$2 | Multiply by the number of tables being actively written to. With 10 tables at 100ms intervals, that's ~$1,400/month in API charges alone. With NFS, that same @@ -212,7 +212,7 @@ Tiering requires files over 128 KiB. ### Throttle window ```ini -replication.primary.throttle.window.duration=10000 # 10 seconds (default) +replication.primary.throttle.window.duration=1000 # 1 second (default) ``` Maximum time before uploading an incomplete segment. If a segment hasn't reached @@ -223,8 +223,8 @@ segments fill up before upload, reducing redundant uploads (write amplification) |-------|----------| | `50` (50ms) | Ultra-low latency. Best with NFS transport. | | `100` (100ms) | Low latency. Good balance for NFS transport. | -| `1000` (1s) | Low latency for object store transport. | -| `10000` (10s) | Default. Balanced. | +| `1000` (1s) | Default. Low latency for object store transport. | +| `10000` (10s) | 10 second delay OK. Fewer uploads. | | `60000` (60s) | 1 minute delay OK. Fewer uploads. | | `300000` (5 min) | Cost-sensitive. Batches more data. | From a268159e536e66a883cd4290e4937dc9f6b5fce0 Mon Sep 17 00:00:00 2001 From: glasstiger Date: Fri, 2 Oct 2026 22:02:38 +0100 Subject: [PATCH 25/25] docs(nodejs): fix failover, shutdown, and error guidance from review - Full example: raise failover_max_attempts together with the failover time budget; the default cap of 8 attempts ends query failover within seconds, so the minute-long budget alone did nothing - Query failover: state that against refused connections the attempt cap, not the time budget, usually ends failover - Server shutdown: with failover on (the default), an interrupted query fails over like any lost connection; only failover=off reports CANCELLED, which query.cancel() also produces, so do not retry a query you cancelled - Connection errors: document the QwpReconnectExhaustedError wrapper when the first connection is retried, list the cause shapes as bullets, and add an example that unwraps the cause chain - timestampColumn(): show the "us" default unit and extend the 1970 warning to it - Backpressure: a borrowed sender's close() rejects with the append timeout error, a standalone Sender's close() with QwpSenderCloseTimeoutError after close_flush_timeout_millis - Name QwpIngressNackError, which waitForAcknowledged() rejects with for a terminally rejected batch --- documentation/connect/clients/nodejs.md | 137 ++++++++++++++++++------ 1 file changed, 103 insertions(+), 34 deletions(-) diff --git a/documentation/connect/clients/nodejs.md b/documentation/connect/clients/nodejs.md index 2a952460ac..0149c5f89f 100644 --- a/documentation/connect/clients/nodejs.md +++ b/documentation/connect/clients/nodejs.md @@ -298,7 +298,7 @@ The pooled QWP sender and `connectQwpNodeSender()` expose these methods: | `booleanColumn`, `byteColumn`, `shortColumn`, `int32Column` | BOOLEAN, BYTE, SHORT, INT | | `longColumn`, `intColumn` | LONG; safe integer `number` or `bigint` | | `float32Column`, `doubleColumn`, `floatColumn` | FLOAT, DOUBLE, DOUBLE | -| `timestampColumn(name, value, unit?)`, `dateColumn` | TIMESTAMP/TIMESTAMP_NS and DATE | +| `timestampColumn(name, value, unit = "us")`, `dateColumn` | TIMESTAMP/TIMESTAMP_NS and DATE; for units, see [Designated timestamp](#designated-timestamp) | | `charColumn`, `binaryColumn`, `uuidColumn` | CHAR, BINARY (`Uint8Array`), UUID | | `long256Column`, `ipv4Column`, `geohashColumn` | LONG256, IPv4, GEOHASH | | `decimalColumnText`, `decimalColumn`, `decimal64Column`, `decimal128Column`, `decimal256Column` | DECIMAL; see [Decimals](#decimals) | @@ -327,7 +327,10 @@ dictionary full. `at(value, unit)` accepts `"us"` (the default), `"ms"`, or `"ns"`. `Date.now()` is **milliseconds**, so use `.at(Date.now(), "ms")`; -without the unit the row lands in 1970. Nanoseconds require a `bigint`. +without the unit the row lands in 1970. `timestampColumn(name, value, unit)` +takes the same units with the same `"us"` default, so pass the unit there +too: `.timestampColumn("exchange_ts", Date.now(), "ms")`. Nanoseconds require +a `bigint`. `atNow()` asks QuestDB to assign arrival time, which changes on replay. Use the event's timestamp for deduplication. For a newly created table, `"ns"` creates a TIMESTAMP_NS designated timestamp; the other units create @@ -470,9 +473,14 @@ targets 10 GiB with it. When full, publishing waits up to 30 seconds by default, then rejects with `QwpMemoryReplayAppendTimeoutError` or `QwpReplayStoreAppendTimeoutError`. The batch stays staged: slow down and retry `flush()`; do not write the rows again. If you close the sender -instead, `close()` tries once more and, if there is still no room, drops the -staged rows with a warning and rejects with the same error. Configure the cap -with `sf_max_total_bytes` and the wait with `sf_append_deadline_millis`. +instead, it tries once more and, if there is still no room, drops the staged +rows with a warning. A borrowed sender's `close()` then rejects with the same +error, after up to `sf_append_deadline_millis` plus +`close_flush_timeout_millis` (35 seconds by default). A standalone `Sender`'s +`close()` waits at most `close_flush_timeout_millis` (5 seconds by default) +and, with the default deadlines, rejects with `QwpSenderCloseTimeoutError`. +Configure the cap with `sf_max_total_bytes` and the wait with +`sf_append_deadline_millis`. #### Batch size limits @@ -492,9 +500,11 @@ After `flush()`, wait for the cumulative watermark: `await sender.waitForAcknowledged(sender.publishedSequence, 10_000)`. `publishedSequence` includes batches sent by auto-flush; `acknowledgedSequence` -is the last accepted one. `waitForAcknowledged()` rejects on a server -rejection, or with `QwpIngressAckTimeoutError` on timeout; without a timeout -argument, it waits up to `ackTimeoutMs` (15 seconds by default). A timeout +is the last accepted one. `waitForAcknowledged()` rejects with +`QwpIngressNackError` when QuestDB terminally rejects a batch (see +[Ingestion errors](#ingestion-errors)), or with `QwpIngressAckTimeoutError` on +timeout; without a timeout argument, it waits up to `ackTimeoutMs` +(15 seconds by default). A timeout alone does not mean the batch was rejected: it may still be in flight. **Do not use the return value of `flushAndGetSequence()` as the watermark for all your rows**: it returns `-1n` if an earlier auto-flush already published them. @@ -1032,7 +1042,8 @@ Set `timeoutMs` per query (or `egressSession.queryTimeoutMs` by default). A deadline cancels the query and reports `QwpEgressQueryTimeoutError`. Leaving a `for await` loop early also cancels the query, and `completion` then rejects with `QwpEgressQueryAbandonedError`. `query.cancel()` requests -cancellation but does not wait for it. +cancellation but does not wait for it; the query then fails with +`QwpEgressQueryError` and `status` `QWP_STATUS.CANCELLED`. Cancellation is prompt only with a credit window (see [Flow control](#flow-control)). Without one, the server keeps streaming after @@ -1143,8 +1154,8 @@ can import it for `instanceof` checks: | Local value validation | Fix the value; the row in progress was discarded. Test `QwpBatchTooLargeError` before `RangeError` because it extends `RangeError`. | | `QwpMemoryReplayAppendTimeoutError` / `QwpReplayStoreAppendTimeoutError` | The batch stays staged. Slow down and retry the flush, not the rows. | | ACK timeout: `QwpIngressAckTimeoutError` from `waitForAcknowledged()`, or a plain `Error` from a `flush()` that waits for an acknowledgement | Not a rejection: QuestDB may still acknowledge the batch. Do not write the rows again; wait again, or raise `ackTimeoutMs` or `durableAckTimeoutMs`. See [Awaiting acknowledgements](#awaiting-acknowledgements). | -| Server rejection | See [Ingestion errors](#ingestion-errors); a terminal rejection fails the sender. | -| `QwpEgressQueryError` | Check `status`: fix SQL or bind values for `PARSE_ERROR`; retry on a new lease for `CANCELLED`, such as a query cancelled by a server shutdown. The lease remains usable after a SQL error. | +| Server rejection: `QwpIngressNackError` from `waitForAcknowledged()`, or an `onSenderError` report | See [Ingestion errors](#ingestion-errors); a terminal rejection fails the sender. | +| `QwpEgressQueryError` | Check `status`: fix SQL or bind values for `PARSE_ERROR`. `CANCELLED` comes from your own `query.cancel()` or, with `failover=off`, from a server shutdown: retry on a new lease only a query you did not cancel. The lease remains usable after a SQL error. | | `QwpEgressSessionClosedError` | The query connection was lost with failover off. Close the lease and retry the whole query. | | `QwpReconnectExhaustedError` | Close the failed sender or query lease and borrow a new one. | @@ -1220,8 +1231,9 @@ category has a fixed default policy, because Node.js does not apply the The handler also runs for retriable rejections, which the client resends: only `terminal` (the sender stops) and `abandoned` (journal data set aside) mean the rows are not being delivered. Without a callback, the client logs -rejections. `waitForAcknowledged()` -rejects for a rejected batch. Branch on category, not the unstable message +rejections. `waitForAcknowledged()` rejects with `QwpIngressNackError` for a +terminally rejected batch; its `senderError` property has the fields above. +Branch on category, not the unstable message text, and redact messages before sending them to external trackers. `ingressSession.onError` receives an event with `error`, `terminal`, `timestampMs`, and, for a server rejection, `senderError`. It also reports @@ -1309,9 +1321,14 @@ try { ``` `requestId` numbers queries per connection; it is not a server-side -correlation ID. A server that shuts down cancels running queries: they fail -with `status` `QWP_STATUS.CANCELLED` instead of failing over, so retry them on -a new lease. Other failures are separate classes, not `QwpEgressQueryError`: +correlation ID. A server that shuts down interrupts its running queries. With +failover on (the default), the client treats this as a lost connection: it +runs the query again from its first row, or fails with +`QwpReconnectExhaustedError` once failover gives up; see +[Query failover](#query-failover). With `failover=off`, the query fails with +`status` `QWP_STATUS.CANCELLED`; retry it on a new lease. Your own +`query.cancel()` also produces `CANCELLED`, so do not retry a query you +cancelled. Other failures are separate classes, not `QwpEgressQueryError`: - `QwpEgressQueryTimeoutError`: `timeoutMs` expired. - `QwpEgressQueryAbandonedError`: the loop ended early; see @@ -1325,20 +1342,68 @@ another. ### Connection-level errors -The pool wraps connection creation failures in `QwpPoolResourceError`; inspect -its `cause`. A `QwpUpgradeError` covers any failure while opening the -WebSocket, so check its `kind`: `authentication` for an HTTP `401` or `403`, -`transport` or `timeout` when QuestDB is unreachable, and others such as -`role-rejected` and `version-mismatch`. `QwpRoleMismatchError` and -`QwpDurableAckUnavailableError` extend `QwpUpgradeError`, so test for them -first. With several endpoints, a connection that fails on every endpoint, for -example because none is reachable, has a `QwpFailoverError` cause whose -`attempts` records each endpoint's failure. An authentication rejection is the -exception: it stops the endpoint sweep at once, so the cause is that -`QwpUpgradeError` itself. With one endpoint, the cause is always that -endpoint's error. A borrow at pool capacity times out as -`QwpPoolAcquireTimeoutError` after `acquire_timeout_ms` (5 seconds by -default). +The pool wraps a failure to open a connection in `QwpPoolResourceError`. +Follow its `cause` chain to the reason: + +- **Retried first connection.** If the first connection is retried, the next + cause is a `QwpReconnectExhaustedError` that wraps the last attempt's error. + Queries retry it with `failover=on`, any `failover_*` key, or a typed + `egressSession.reconnect`; senders retry it with `initial_connect_retry=on` + or any `reconnect_*` key (see [Startup and outage modes](#ingestion-modes)). +- **Several endpoints.** When every endpoint fails, for example because none is + reachable, the error is a `QwpFailoverError` whose `attempts` records each + endpoint's failure. +- **One endpoint.** The error is that endpoint's `QwpUpgradeError`. +- **Authentication.** A `401` or `403` stops the endpoint sweep and is not + retried, so the cause is the `QwpUpgradeError` itself, even with several + endpoints or a retried first connection. + +A `QwpUpgradeError` covers any failure while opening the WebSocket, so check +its `kind`: `authentication` for an HTTP `401` or `403`, `transport` or +`timeout` when QuestDB is unreachable, and others such as `role-rejected` and +`version-mismatch`. `QwpRoleMismatchError` and `QwpDurableAckUnavailableError` +extend `QwpUpgradeError`, so test for them first. A borrow at pool capacity +times out as `QwpPoolAcquireTimeoutError` after `acquire_timeout_ms` +(5 seconds by default). + +```typescript +import { + connectQwpNodeClient, + QwpFailoverError, + QwpPoolResourceError, + QwpReconnectExhaustedError, + QwpUpgradeError, +} from "@questdb/nodejs-client"; + +// Skip the pool and retry wrappers to reach the connection error itself. +function connectionError(error: unknown): unknown { + let cause = error; + while ( + (cause instanceof QwpPoolResourceError || + cause instanceof QwpReconnectExhaustedError) && + cause.cause !== undefined + ) { + cause = cause.cause; + } + return cause; +} + +try { + const db = await connectQwpNodeClient("ws::addr=localhost:9000;"); + console.log("connected to QuestDB"); + await db.close(); +} catch (error) { + const cause = connectionError(error); + if (cause instanceof QwpUpgradeError && cause.kind === "authentication") { + console.error("QuestDB rejected the credentials"); + } else if (cause instanceof QwpFailoverError) { + console.error("no endpoint accepted the connection:", cause.attempts); + } else { + console.error("cannot connect to QuestDB:", cause); + } + process.exitCode = 1; +} +``` An authentication rejection never moves the client to another endpoint. It is terminal before a sender's first successful connection. After that, senders @@ -1471,8 +1536,10 @@ Between failover attempts, the client waits a random delay below a ceiling that starts at `failover_backoff_initial_ms` (50 ms by default) and doubles up to `failover_backoff_max_ms` (1 second). It makes at most `failover_max_attempts` (8) attempts within `failover_max_duration_ms` -(30 seconds), so the attempt limit can end failover before the time budget: -raise `failover_max_attempts` for longer outages. When failover gives up, the +(30 seconds). Against refused connections, as while a server restarts, eight +attempts take only a few seconds, so the attempt limit ends failover long +before the time budget: raise `failover_max_attempts` together with +`failover_max_duration_ms` for longer outages. When failover gives up, the query fails with `QwpReconnectExhaustedError`; close the lease and borrow a new one. @@ -1739,8 +1806,10 @@ const db = await connectQwpNodeClient( // Start while QuestDB is down and journal rows across restarts. "lazy_connect=on;sf_dir=/var/lib/myapp/qdb-sf;sender_id=trades;" + "sf_max_segment_bytes=1m;" + - // Bound query buffering; let query failover run for up to a minute. - "initial_credit=1048576;failover_max_duration_ms=60000;", + // Bound query buffering. Let query failover retry for up to a minute: + // the default cap of 8 attempts would otherwise end it within seconds. + "initial_credit=1048576;failover_max_attempts=1000;" + + "failover_max_duration_ms=60000;", { ingressSession: { onSenderError: (error) => {