From 4a52d446c46449469f643970b9affd90abb057a9 Mon Sep 17 00:00:00 2001 From: MisterRaindrop <278811821@qq.com> Date: Wed, 5 Aug 2026 11:32:13 +0800 Subject: [PATCH] interconnect: restart in fast mode so the test does not race crash recovery The test sets shared_preload_libraries and restarts with "gpstop -raiq". An immediate shutdown skips the shutdown checkpoint, so the control file is left in a state other than DB_SHUTDOWNED and the next startup performs crash recovery: xlogrecovery.c sets InRecovery, xlog.c calls PerformWalRecovery(), which signals PMSIGNAL_RECOVERY_STARTED, and the postmaster moves to PM_RECOVERY. In that state canAcceptConnections() answers CAC_NOTCONSISTENT, reported as "the database system is not accepting connections" with detail "Hot standby mode is disabled". gpstart makes exactly such a connection right after pg_ctl returns, to read the segment configuration, so gpstop -r exits CRITICAL and the restart is reported as failed. The damage does not stop there. psql gives up at the \c that follows, so every statement in the file is skipped and the test fails as a whole; the cleanup at the end of the file never runs; and gpstart never got past starting the coordinator in admin mode, so the cluster is left with no segments up. Suites that run after this one in the same job then lose their Gather Motion nodes and fail as well. Shut down fast instead. A fast shutdown writes the shutdown checkpoint, the control file says DB_SHUTDOWNED, no recovery runs, PM_RECOVERY is never entered, and CAC_NOTCONSISTENT cannot be returned -- the failure becomes unreachable rather than merely less likely. Fast is also what the rest of the tree already uses: gpstop -raf/-arf appear in dozens of places, and this file was the only user of -raiq. Measured on a three-segment demo cluster, dirtying 1.5M coordinator rows before each restart so that recovery is slow enough to lose the race reliably: -raiq failed 2/2 with the message above, -rafq passed 3/3 with all three segments still up afterwards. pg_controldata confirms the mechanism at the other end -- "in production" after an immediate shutdown, "shut down" after a fast one. The test still passes under pg_regress with the change. --- contrib/interconnect/sql/interconnect.sql | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/contrib/interconnect/sql/interconnect.sql b/contrib/interconnect/sql/interconnect.sql index 4e6555b6b82..32bbef92295 100644 --- a/contrib/interconnect/sql/interconnect.sql +++ b/contrib/interconnect/sql/interconnect.sql @@ -1,6 +1,12 @@ -- start_ignore \! gpconfig -c shared_preload_libraries -v "interconnect" -\! gpstop -raiq +-- Restart in fast mode, not immediate: an immediate shutdown skips the +-- shutdown checkpoint, so the next startup runs crash recovery, and while the +-- postmaster is in PM_RECOVERY it rejects the connection gpstart makes to read +-- the segment configuration ("the database system is not accepting +-- connections"). gpstop -r then fails, psql gives up at the \c below, and the +-- whole file is skipped with only the coordinator left running. +\! gpstop -rafq \c DROP TABLE IF EXISTS test_ic_data; CREATE EXTENSION IF NOT EXISTS interconnect; @@ -83,5 +89,5 @@ DROP EXTENSION interconnect; -- start_ignore \! gpconfig -r shared_preload_libraries -\! gpstop -raiq +\! gpstop -rafq -- end_ignore