Skip to content

SuperServer: no attachment to ANY database can proceed while two or more attachments wait for a database under exclusive validation #9169

Description

@fdcastel

While a database is being validated exclusively (Services API isc_action_svc_repair with isc_spb_rpr_validate_db | isc_spb_rpr_full, i.e. gfix -v -full), attachments to that database wait for the validation to finish, which is expected. But as soon as a second attachment is waiting for the validated database, every new attachment to every other database on the same SuperServer hangs too, until the validation ends. One waiting attachment is harmless; the second one takes the whole server down for new connections. Existing attachments keep working. Clients see the attach hang for the duration of the validation (or connection rejected by remote interface once ConnectionTimeout expires — 180 s by default).

Reproduced on Windows x64 with the official zip kits (SuperServer, stock firebird.conf plus the port), TCP and XNET alike, on every version I tried:

Version matrix

Two databases on the same server: BIG (12 GB, or 4 GB where noted) and SMALL (25 MB). Exclusive full validation of BIG starts at T0; K "waiters" attach to BIG (one isql process each, staggered 6 s); a probe attaches to SMALL every 2 s and runs SELECT 1 FROM RDB$DATABASE. Normal probe: 150-250 ms (process start included).

Engine Waiters Validation Probes of SMALL Blocked from
WI-V3.0.14.33856 1 22 s 16/16 OK, max 191 ms —
WI-V3.0.14.33856 2 22 s 4 of 16 hung (max 10.7 s) 2nd waiter → end
WI-V4.0.5.3140 1 109 s 59/59 OK, max 220 ms —
WI-V4.0.5.3140 2 122 s 53 of 66 hung (max 110.8 s) 2nd waiter → end
WI-V4.0.7.3271 1 121 s 65/65 OK, max 242 ms —
WI-V4.0.7.3271 2 100 s 43 of 55 hung (max 89.7 s) 2nd waiter → end
WI-V5.0.3.1683 1 102 s 56/56 OK, max 176 ms —
WI-V5.0.3.1683 2 105 s 44 of 57 hung (max 93.6 s) 2nd waiter → end
WI-V5.0.4.1812 0 6 s (4 GB) 8/8 OK —
WI-V5.0.4.1812 1 54 s (4 GB) 32/32 OK, max 193 ms —
WI-V5.0.4.1812 2 45 s (4 GB) 24 of 52 hung (max 37.2 s) 2nd waiter → end
WI-V5.0.4.1812 3 38 s (4 GB) 13 of 24 hung (max 31.4 s) 2nd waiter → end

"Blocked from 2nd waiter": in every run the first hung probe is the first one started after the second waiter's isql was launched (within 50 ms), and every probe started until the validation's end hangs; they are all released within a second of the validation finishing. Probes started while only one waiter existed (up to 8 s of them) were normal.

Variants on 5.0.4 (4 GB BIG), all with 2 waiters:

Variant Result
waiters over XNET, probes over TCP same: hung from the 2nd waiter (21 of 33)
waiters over TCP, probes over XNET same: hung from the 2nd waiter (17 of 28)
both waiter processes killed 6 s after the 2nd one attached no change: probes stayed hung until the validation ended (18 of 28); the server keeps the attachments of the dead clients
server ConnectionTimeout = 20 (see below)

Reproduction

repro-attach-freeze.ps1 (at the end of this issue) does all of this against any running server: creates BIG and SMALL, starts the validation with fbsvcmgr, starts the waiters, prints one line per probe. By hand it is five commands:

# 1. two databases (BIG large enough for the validation to take a minute; 4 GB of 1 KB rows is enough)
isql -q -user SYSDBA -password masterkey    # CREATE DATABASE 'localhost:C:\t\big.fdb' ...; CREATE DATABASE 'localhost:C:\t\small.fdb' ...

# 2. T0: exclusive full validation through the Services API (what gfix -v -full does)
fbsvcmgr localhost:service_mgr -user SYSDBA -password masterkey -action_repair -dbname C:\t\big.fdb -rpr_validate_db -rpr_full

# 3. while it runs: ONE attachment to BIG waits (expected)
isql -user SYSDBA -password masterkey localhost:C:\t\big.fdb        # hangs until the validation ends

# 4. attach to SMALL: still fine
isql -user SYSDBA -password masterkey localhost:C:\t\small.fdb      # answers at once

# 5. a SECOND attachment to BIG waits...
isql -user SYSDBA -password masterkey localhost:C:\t\big.fdb        # hangs

# 6. ...and now SMALL (and any other database) is unreachable until the validation ends
isql -user SYSDBA -password masterkey localhost:C:\t\small.fdb      # hangs

What the threads are doing (full dump of firebird.exe 5.0.4.1812 during the hang, with the release PDBs)

Four kinds of threads, forming one chain:

  1. The validator (service thread): Validation::run → walk_database → walk_relation → walk_pointer_page → walk_data_page → CCH_fetch_page → PIO_read, inside JProvider::internalAttach (the isc_dpb_verify attach).

  2. Waiter 1 (worker thread, attach_database → JProvider::internalAttach → PAG_header → CCH_FETCH → CCH_fetch_lock → get_buffer → BufferDesc::addRef → SyncObject::lock → wait): blocked on the header page latch. Offline validation keeps the header page latched for its whole run — Validation::walk_database fetches it with fetch_page(true, HEADER_PAGE, …) and releases it only if (vdr_flags & VDR_online) (validation.cpp 1617-1621 in v5.0.4). PAG_header is called from internalAttach before initGuard.leave() ("Basic DBB initialization complete", jrd.cpp ~1899), so this thread is blocked while holding the database's dbb_init_fini mutex.

  3. Waiter 2 (worker thread, attach_database → JProvider::internalAttach → initAttachment → RefMutexUnlock::enter → Database::ExistenceRefMutex::enter → Mutex::enter): blocked on that dbb_init_fini. initAttachment runs with the process-global dbInitMutex held (guardDbInit.enter() at jrd.cpp 1696 on Windows / 1742 elsewhere, released at 1760 only after initAttachment returns for an existing database), so this thread is blocked while holding dbInitMutex.

  4. Every other attachment (5 worker threads in the dump, all probes of SMALL: attach_database → JProvider::internalAttach → EnsureUnlock<Mutex>::enter → Mutex::enter): blocked on dbInitMutex at guardDbInit.enter().

So: header page latch (validator) → dbb_init_fini of BIG (waiter 1) → dbInitMutex (waiter 2) → everyone. With a single waiter the chain stops at the per-database mutex and nothing else is affected, which is why one waiter is harmless and two are not. The same code (the guardDbInit/initGuard ordering around PAG_header, and the header page held by offline validation) is present in 3.0.14 and 4.0.7, matching the matrix.

Killing the waiting clients does not help: the worker threads are inside the engine and never look at the socket; the server keeps their attachments (and the chain) until the validation ends.

What does and does not change it

Result
stock firebird.conf vs DefaultDbCachePages = 128K + FileSystemCacheThreshold = 100M same
TCP vs XNET (waiters or probes) same
1 waiter never reproduces (tested up to 120 s windows on every version)
2 or 3 waiters always reproduces, from the second waiter's attach to the validation's end
killing the waiters' processes no change
ConnectionTimeout = 20 (server and/or client), probes every 2 s no rejection: hung attaches waited up to 100 s and then succeeded
client ConnectionTimeout = 20, probes every 30 s every probe fails after exactly 20 s with SQLSTATE = 08004 connection rejected by remote interface; the server keeps its socket in CLOSE_WAIT and its worker thread blocked until the validation ends

Two ways the client sees the same freeze, depending on where its connection blocks:

  • Connection made within ~10 s of another one: the Srp server plugin still has its cached security-database attachment (SecDbCache.cpp: released by a 10 s timer), authentication completes, the client gets op_accept, sends op_attach, and its worker blocks on dbInitMutex inside the target database's attach. No timeout applies after op_accept: the client hangs until the validation ends.
  • Connection made after a quiet period: the security-database attachment has been closed, so authenticating the new connection needs a fresh attach to the security database — which goes through the same JProvider::internalAttach and blocks on dbInitMutex before op_accept is sent. The client is still in the connect phase, its ConnectionTimeout (default 180 s) expires, and it reports connection rejected by remote interface. The server's worker stays blocked, so every such client leaves one blocked thread and one CLOSE_WAIT socket behind until the validation ends. With sparse connections this is what one sees: every client failing after 180 s while blocked threads pile up on the server.

Environment

  • Windows 11 x64; official ZIP kits run as applications (firebird -a), ServerMode = Super, one instance per version on its own port.
  • Databases created by the same kit (page size 8192, ODS 12 / 13.0 / 13.1 respectively); 12 GB for 3.0/4.0/5.0.3, 4 GB for 5.0.4 (where the validation was slowed by a concurrent disk load; the window is what matters).
  • Reproduced on 3.0.14.33856, 4.0.5.3140, 4.0.7.3271, 5.0.3.1683, 5.0.4.1812.

repro-attach-freeze.ps1

PowerShell 7 script (creates the two databases, runs the validation, the waiters and the probe against any running server)
#Requires -Version 7
<#
.SYNOPSIS
    Self-contained reproduction: Firebird SuperServer stops accepting attachments to EVERY database
    while two or more attachments wait for a database that is under exclusive validation.

    Needs a running Firebird server (3.0, 4.0 or 5.0, ServerMode = Super) reachable over TCP, and the
    kit's isql.exe / fbsvcmgr.exe. Nothing is installed. Creates two databases under -WorkDir:
    BIG (-BigMB, default 8192) and SMALL (~25 MB). Then:

      T0       exclusive full validation of BIG through the Services API
               (fbsvcmgr -action_repair -rpr_validate_db -rpr_full = gfix -v -full)
      T0+2s    waiter 1: one isql attaches to BIG (blocks until the validation ends)
      T0+5s    waiter 2: a second isql attaches to BIG
      always   every 2 s one isql attaches to SMALL and runs SELECT 1 FROM RDB$DATABASE; its duration
               is printed. Expected: a few ms. Observed: from waiter 2 on, every attach to SMALL hangs
               until the validation finishes.

    Run with -Waiters 1 to see that a single waiter is harmless, -Waiters 0 for the control.

.EXAMPLE
    .\repro-attach-freeze.ps1 -FirebirdDir 'C:\Program Files\Firebird\Firebird_5_0' -Server 127.0.0.1 -Port 3050 -Waiters 2
#>
param(
    [Parameter(Mandatory)][string] $FirebirdDir,     # folder with isql.exe and fbsvcmgr.exe
    [string] $Server = '127.0.0.1',
    [int] $Port = 3050,
    [string] $User = 'SYSDBA',
    [string] $Password = 'masterkey',
    [string] $WorkDir = (Join-Path $PWD 'attach-freeze'),
    [int] $BigMB = 8192,          # the validation must outlast the waiters' arrival; raise it if the script says so
    [int] $Waiters = 2,
    [int] $WaiterDelaySeconds = 2,
    [int] $WaiterStaggerSeconds = 3,
    [int] $ProbeIntervalSeconds = 2
)
$ErrorActionPreference = 'Stop'
[cultureinfo]::CurrentCulture = [cultureinfo]::InvariantCulture
$isql = Join-Path $FirebirdDir 'isql.exe'; $svc = Join-Path $FirebirdDir 'fbsvcmgr.exe'
foreach ($t in $isql, $svc) { if (-not (Test-Path $t)) { throw "not found: $t" } }
New-Item -ItemType Directory -Path $WorkDir -Force | Out-Null
$big = Join-Path $WorkDir 'big.fdb'; $small = Join-Path $WorkDir 'small.fdb'
$tcp = "$Server/${Port}:"

function Invoke-Sql([string] $Sql, [string] $Db = '') {
    $f = Join-Path $WorkDir ([guid]::NewGuid().ToString('N') + '.sql'); Set-Content $f $Sql -Encoding ascii
    try { $out = & $isql -q -b -user $User -password $Password -i $f $Db 2>&1; if ($LASTEXITCODE) { throw "isql failed: $(($out | ForEach-Object { "$_" }) -join ' ')" } } finally { Remove-Item $f }
}
function New-Db([string] $Path, [int] $TargetMB, [int] $RowsPerBatch = 100000) {
    if (Test-Path $Path) { Write-Host "reusing $Path"; return }
    Invoke-Sql "CREATE DATABASE '$tcp$Path' USER '$User' PASSWORD '$Password' PAGE_SIZE 8192;`nCREATE TABLE T (ID BIGINT NOT NULL PRIMARY KEY, K INTEGER NOT NULL, S VARCHAR(1000) NOT NULL);`nCREATE INDEX IX_T_K ON T (K);`nCOMMIT;`nQUIT;"
    $batch = 0
    while ((Get-Item $Path).Length / 1MB -lt $TargetMB) {
        $from = $batch * $RowsPerBatch; $batch++
        Invoke-Sql -Db "$tcp$Path" -Sql @"
SET TERM ^;
EXECUTE BLOCK AS
  DECLARE i BIGINT = $from;
  DECLARE s VARCHAR(1300);
BEGIN
  WHILE (i < $from + $RowsPerBatch) DO
  BEGIN
    s = uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid());
    s = s || s || s || s || s || s || s;
    INSERT INTO T (ID, K, S) VALUES (:i, MOD(:i * 7919, 1000003), SUBSTRING(:s FROM 1 FOR 980));
    i = i + 1;
  END
END^
SET TERM ;^
COMMIT;
QUIT;
"@
        Write-Host ("  {0}: {1} MB" -f (Split-Path $Path -Leaf), [int]((Get-Item $Path).Length / 1MB))
    }
}
Write-Host "creating databases under $WorkDir (BIG $BigMB MB, SMALL 25 MB)..."
New-Db $small 25 20000
New-Db $big $BigMB

$probeSql = Join-Path $WorkDir 'probe.sql'; "SELECT 1 FROM RDB`$DATABASE;`nQUIT;`n" | Set-Content $probeSql -Encoding ascii
$probe = { param($Isql, $Sql, $Target, $User, $Password, $T0)
    $started = Get-Date; $sw = [Diagnostics.Stopwatch]::StartNew()
    $out = & $Isql -q -user $User -password $Password -i $Sql $Target 2>&1
    '{0,8:n1}s  probe SMALL  {1,7} ms  exit {2}  {3}' -f ($started - $T0).TotalSeconds, $sw.ElapsedMilliseconds, $LASTEXITCODE, ((($out | ForEach-Object { "$_" }) -join ' ') -replace '\s+', ' ').Trim()
}

Write-Host ''
Write-Host ("T0: exclusive validation of BIG ({0} MB) starts; {1} waiter(s) at +{2}s then every {3}s; probe of SMALL every {4}s" -f [int]((Get-Item $big).Length / 1MB), $Waiters, $WaiterDelaySeconds, $WaiterStaggerSeconds, $ProbeIntervalSeconds)
$T0 = Get-Date
$val = Start-Process -FilePath $svc -ArgumentList @("${tcp}service_mgr", '-user', $User, '-password', $Password, '-action_repair', '-dbname', $big, '-rpr_validate_db', '-rpr_full') -NoNewWindow -PassThru
$waiterProcs = @(); $nextWaiter = $T0.AddSeconds($WaiterDelaySeconds); $left = $Waiters; $jobs = @()
while (-not $val.HasExited) {
    $now = Get-Date
    if ($left -gt 0 -and $now -ge $nextWaiter) {
        $waiterProcs += Start-Process -FilePath $isql -ArgumentList @('-q', '-user', $User, '-password', $Password, '-i', $probeSql, "$tcp$big") -NoNewWindow -PassThru -RedirectStandardOutput (Join-Path $WorkDir "waiter-$($Waiters - $left + 1).out")
        Write-Host ('{0,8:n1}s  waiter {1} attaches to BIG' -f ($now - $T0).TotalSeconds, ($Waiters - $left + 1))
        $left--; $nextWaiter = $now.AddSeconds($WaiterStaggerSeconds)
    }
    $jobs += Start-ThreadJob -ScriptBlock $probe -ArgumentList $isql, $probeSql, "$tcp$small", $User, $Password, $T0 -ThrottleLimit 128
    foreach ($j in @($jobs | Where-Object State -eq 'Completed')) { Receive-Job $j | Write-Host; Remove-Job $j; $jobs = @($jobs | Where-Object Id -ne $j.Id) }
    Start-Sleep -Seconds $ProbeIntervalSeconds
}
Write-Host ('{0,8:n1}s  validation finished (exit {1})' -f ((Get-Date) - $T0).TotalSeconds, $val.ExitCode)
if ($left -gt 0) { Write-Warning "the validation ended before all $Waiters waiters had attached: nothing was tested. Raise -BigMB (the validation must last well past +$($WaiterDelaySeconds + ($Waiters - 1) * $WaiterStaggerSeconds) s)." }
$deadline = (Get-Date).AddSeconds(200)
while ($jobs -and (Get-Date) -lt $deadline) { foreach ($j in @($jobs | Where-Object State -ne 'Running')) { Receive-Job $j | Write-Host; Remove-Job $j; $jobs = @($jobs | Where-Object Id -ne $j.Id) }; Start-Sleep -Seconds 1 }
$waiterProcs | Where-Object { -not $_.HasExited } | ForEach-Object { Stop-Process -Id $_.Id -Force }

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions