While a database is being validated exclusively (Services API isc_action_svc_repair with isc_spb_rpr_validate_db | isc_spb_rpr_full, i.e. gfix -v -full), attachments to that database wait for the validation to finish, which is expected. But as soon as a second attachment is waiting for the validated database, every new attachment to every other database on the same SuperServer hangs too, until the validation ends. One waiting attachment is harmless; the second one takes the whole server down for new connections. Existing attachments keep working. Clients see the attach hang for the duration of the validation (or connection rejected by remote interface once ConnectionTimeout expires — 180 s by default).
Reproduced on Windows x64 with the official zip kits (SuperServer, stock firebird.conf plus the port), TCP and XNET alike, on every version I tried:
Version matrix
Two databases on the same server: BIG (12 GB, or 4 GB where noted) and SMALL (25 MB). Exclusive full validation of BIG starts at T0; K "waiters" attach to BIG (one isql process each, staggered 6 s); a probe attaches to SMALL every 2 s and runs SELECT 1 FROM RDB$DATABASE. Normal probe: 150-250 ms (process start included).
| Engine |
Waiters |
Validation |
Probes of SMALL |
Blocked from |
| WI-V3.0.14.33856 |
1 |
22 s |
16/16 OK, max 191 ms |
— |
| WI-V3.0.14.33856 |
2 |
22 s |
4 of 16 hung (max 10.7 s) |
2nd waiter → end |
| WI-V4.0.5.3140 |
1 |
109 s |
59/59 OK, max 220 ms |
— |
| WI-V4.0.5.3140 |
2 |
122 s |
53 of 66 hung (max 110.8 s) |
2nd waiter → end |
| WI-V4.0.7.3271 |
1 |
121 s |
65/65 OK, max 242 ms |
— |
| WI-V4.0.7.3271 |
2 |
100 s |
43 of 55 hung (max 89.7 s) |
2nd waiter → end |
| WI-V5.0.3.1683 |
1 |
102 s |
56/56 OK, max 176 ms |
— |
| WI-V5.0.3.1683 |
2 |
105 s |
44 of 57 hung (max 93.6 s) |
2nd waiter → end |
| WI-V5.0.4.1812 |
0 |
6 s (4 GB) |
8/8 OK |
— |
| WI-V5.0.4.1812 |
1 |
54 s (4 GB) |
32/32 OK, max 193 ms |
— |
| WI-V5.0.4.1812 |
2 |
45 s (4 GB) |
24 of 52 hung (max 37.2 s) |
2nd waiter → end |
| WI-V5.0.4.1812 |
3 |
38 s (4 GB) |
13 of 24 hung (max 31.4 s) |
2nd waiter → end |
"Blocked from 2nd waiter": in every run the first hung probe is the first one started after the second waiter's isql was launched (within 50 ms), and every probe started until the validation's end hangs; they are all released within a second of the validation finishing. Probes started while only one waiter existed (up to 8 s of them) were normal.
Variants on 5.0.4 (4 GB BIG), all with 2 waiters:
| Variant |
Result |
| waiters over XNET, probes over TCP |
same: hung from the 2nd waiter (21 of 33) |
| waiters over TCP, probes over XNET |
same: hung from the 2nd waiter (17 of 28) |
| both waiter processes killed 6 s after the 2nd one attached |
no change: probes stayed hung until the validation ended (18 of 28); the server keeps the attachments of the dead clients |
server ConnectionTimeout = 20 |
(see below) |
Reproduction
repro-attach-freeze.ps1 (at the end of this issue) does all of this against any running server: creates BIG and SMALL, starts the validation with fbsvcmgr, starts the waiters, prints one line per probe. By hand it is five commands:
# 1. two databases (BIG large enough for the validation to take a minute; 4 GB of 1 KB rows is enough)
isql -q -user SYSDBA -password masterkey # CREATE DATABASE 'localhost:C:\t\big.fdb' ...; CREATE DATABASE 'localhost:C:\t\small.fdb' ...
# 2. T0: exclusive full validation through the Services API (what gfix -v -full does)
fbsvcmgr localhost:service_mgr -user SYSDBA -password masterkey -action_repair -dbname C:\t\big.fdb -rpr_validate_db -rpr_full
# 3. while it runs: ONE attachment to BIG waits (expected)
isql -user SYSDBA -password masterkey localhost:C:\t\big.fdb # hangs until the validation ends
# 4. attach to SMALL: still fine
isql -user SYSDBA -password masterkey localhost:C:\t\small.fdb # answers at once
# 5. a SECOND attachment to BIG waits...
isql -user SYSDBA -password masterkey localhost:C:\t\big.fdb # hangs
# 6. ...and now SMALL (and any other database) is unreachable until the validation ends
isql -user SYSDBA -password masterkey localhost:C:\t\small.fdb # hangs
What the threads are doing (full dump of firebird.exe 5.0.4.1812 during the hang, with the release PDBs)
Four kinds of threads, forming one chain:
-
The validator (service thread): Validation::run → walk_database → walk_relation → walk_pointer_page → walk_data_page → CCH_fetch_page → PIO_read, inside JProvider::internalAttach (the isc_dpb_verify attach).
-
Waiter 1 (worker thread, attach_database → JProvider::internalAttach → PAG_header → CCH_FETCH → CCH_fetch_lock → get_buffer → BufferDesc::addRef → SyncObject::lock → wait): blocked on the header page latch. Offline validation keeps the header page latched for its whole run — Validation::walk_database fetches it with fetch_page(true, HEADER_PAGE, …) and releases it only if (vdr_flags & VDR_online) (validation.cpp 1617-1621 in v5.0.4). PAG_header is called from internalAttach before initGuard.leave() ("Basic DBB initialization complete", jrd.cpp ~1899), so this thread is blocked while holding the database's dbb_init_fini mutex.
-
Waiter 2 (worker thread, attach_database → JProvider::internalAttach → initAttachment → RefMutexUnlock::enter → Database::ExistenceRefMutex::enter → Mutex::enter): blocked on that dbb_init_fini. initAttachment runs with the process-global dbInitMutex held (guardDbInit.enter() at jrd.cpp 1696 on Windows / 1742 elsewhere, released at 1760 only after initAttachment returns for an existing database), so this thread is blocked while holding dbInitMutex.
-
Every other attachment (5 worker threads in the dump, all probes of SMALL: attach_database → JProvider::internalAttach → EnsureUnlock<Mutex>::enter → Mutex::enter): blocked on dbInitMutex at guardDbInit.enter().
So: header page latch (validator) → dbb_init_fini of BIG (waiter 1) → dbInitMutex (waiter 2) → everyone. With a single waiter the chain stops at the per-database mutex and nothing else is affected, which is why one waiter is harmless and two are not. The same code (the guardDbInit/initGuard ordering around PAG_header, and the header page held by offline validation) is present in 3.0.14 and 4.0.7, matching the matrix.
Killing the waiting clients does not help: the worker threads are inside the engine and never look at the socket; the server keeps their attachments (and the chain) until the validation ends.
What does and does not change it
|
Result |
stock firebird.conf vs DefaultDbCachePages = 128K + FileSystemCacheThreshold = 100M |
same |
| TCP vs XNET (waiters or probes) |
same |
| 1 waiter |
never reproduces (tested up to 120 s windows on every version) |
| 2 or 3 waiters |
always reproduces, from the second waiter's attach to the validation's end |
| killing the waiters' processes |
no change |
ConnectionTimeout = 20 (server and/or client), probes every 2 s |
no rejection: hung attaches waited up to 100 s and then succeeded |
client ConnectionTimeout = 20, probes every 30 s |
every probe fails after exactly 20 s with SQLSTATE = 08004 connection rejected by remote interface; the server keeps its socket in CLOSE_WAIT and its worker thread blocked until the validation ends |
Two ways the client sees the same freeze, depending on where its connection blocks:
- Connection made within ~10 s of another one: the Srp server plugin still has its cached security-database attachment (
SecDbCache.cpp: released by a 10 s timer), authentication completes, the client gets op_accept, sends op_attach, and its worker blocks on dbInitMutex inside the target database's attach. No timeout applies after op_accept: the client hangs until the validation ends.
- Connection made after a quiet period: the security-database attachment has been closed, so authenticating the new connection needs a fresh attach to the security database — which goes through the same
JProvider::internalAttach and blocks on dbInitMutex before op_accept is sent. The client is still in the connect phase, its ConnectionTimeout (default 180 s) expires, and it reports connection rejected by remote interface. The server's worker stays blocked, so every such client leaves one blocked thread and one CLOSE_WAIT socket behind until the validation ends. With sparse connections this is what one sees: every client failing after 180 s while blocked threads pile up on the server.
Environment
- Windows 11 x64; official ZIP kits run as applications (
firebird -a), ServerMode = Super, one instance per version on its own port.
- Databases created by the same kit (page size 8192, ODS 12 / 13.0 / 13.1 respectively); 12 GB for 3.0/4.0/5.0.3, 4 GB for 5.0.4 (where the validation was slowed by a concurrent disk load; the window is what matters).
- Reproduced on 3.0.14.33856, 4.0.5.3140, 4.0.7.3271, 5.0.3.1683, 5.0.4.1812.
repro-attach-freeze.ps1
PowerShell 7 script (creates the two databases, runs the validation, the waiters and the probe against any running server)
#Requires -Version 7
<#
.SYNOPSIS
Self-contained reproduction: Firebird SuperServer stops accepting attachments to EVERY database
while two or more attachments wait for a database that is under exclusive validation.
Needs a running Firebird server (3.0, 4.0 or 5.0, ServerMode = Super) reachable over TCP, and the
kit's isql.exe / fbsvcmgr.exe. Nothing is installed. Creates two databases under -WorkDir:
BIG (-BigMB, default 8192) and SMALL (~25 MB). Then:
T0 exclusive full validation of BIG through the Services API
(fbsvcmgr -action_repair -rpr_validate_db -rpr_full = gfix -v -full)
T0+2s waiter 1: one isql attaches to BIG (blocks until the validation ends)
T0+5s waiter 2: a second isql attaches to BIG
always every 2 s one isql attaches to SMALL and runs SELECT 1 FROM RDB$DATABASE; its duration
is printed. Expected: a few ms. Observed: from waiter 2 on, every attach to SMALL hangs
until the validation finishes.
Run with -Waiters 1 to see that a single waiter is harmless, -Waiters 0 for the control.
.EXAMPLE
.\repro-attach-freeze.ps1 -FirebirdDir 'C:\Program Files\Firebird\Firebird_5_0' -Server 127.0.0.1 -Port 3050 -Waiters 2
#>
param(
[Parameter(Mandatory)][string] $FirebirdDir, # folder with isql.exe and fbsvcmgr.exe
[string] $Server = '127.0.0.1',
[int] $Port = 3050,
[string] $User = 'SYSDBA',
[string] $Password = 'masterkey',
[string] $WorkDir = (Join-Path $PWD 'attach-freeze'),
[int] $BigMB = 8192, # the validation must outlast the waiters' arrival; raise it if the script says so
[int] $Waiters = 2,
[int] $WaiterDelaySeconds = 2,
[int] $WaiterStaggerSeconds = 3,
[int] $ProbeIntervalSeconds = 2
)
$ErrorActionPreference = 'Stop'
[cultureinfo]::CurrentCulture = [cultureinfo]::InvariantCulture
$isql = Join-Path $FirebirdDir 'isql.exe'; $svc = Join-Path $FirebirdDir 'fbsvcmgr.exe'
foreach ($t in $isql, $svc) { if (-not (Test-Path $t)) { throw "not found: $t" } }
New-Item -ItemType Directory -Path $WorkDir -Force | Out-Null
$big = Join-Path $WorkDir 'big.fdb'; $small = Join-Path $WorkDir 'small.fdb'
$tcp = "$Server/${Port}:"
function Invoke-Sql([string] $Sql, [string] $Db = '') {
$f = Join-Path $WorkDir ([guid]::NewGuid().ToString('N') + '.sql'); Set-Content $f $Sql -Encoding ascii
try { $out = & $isql -q -b -user $User -password $Password -i $f $Db 2>&1; if ($LASTEXITCODE) { throw "isql failed: $(($out | ForEach-Object { "$_" }) -join ' ')" } } finally { Remove-Item $f }
}
function New-Db([string] $Path, [int] $TargetMB, [int] $RowsPerBatch = 100000) {
if (Test-Path $Path) { Write-Host "reusing $Path"; return }
Invoke-Sql "CREATE DATABASE '$tcp$Path' USER '$User' PASSWORD '$Password' PAGE_SIZE 8192;`nCREATE TABLE T (ID BIGINT NOT NULL PRIMARY KEY, K INTEGER NOT NULL, S VARCHAR(1000) NOT NULL);`nCREATE INDEX IX_T_K ON T (K);`nCOMMIT;`nQUIT;"
$batch = 0
while ((Get-Item $Path).Length / 1MB -lt $TargetMB) {
$from = $batch * $RowsPerBatch; $batch++
Invoke-Sql -Db "$tcp$Path" -Sql @"
SET TERM ^;
EXECUTE BLOCK AS
DECLARE i BIGINT = $from;
DECLARE s VARCHAR(1300);
BEGIN
WHILE (i < $from + $RowsPerBatch) DO
BEGIN
s = uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid()) || uuid_to_char(gen_uuid());
s = s || s || s || s || s || s || s;
INSERT INTO T (ID, K, S) VALUES (:i, MOD(:i * 7919, 1000003), SUBSTRING(:s FROM 1 FOR 980));
i = i + 1;
END
END^
SET TERM ;^
COMMIT;
QUIT;
"@
Write-Host (" {0}: {1} MB" -f (Split-Path $Path -Leaf), [int]((Get-Item $Path).Length / 1MB))
}
}
Write-Host "creating databases under $WorkDir (BIG $BigMB MB, SMALL 25 MB)..."
New-Db $small 25 20000
New-Db $big $BigMB
$probeSql = Join-Path $WorkDir 'probe.sql'; "SELECT 1 FROM RDB`$DATABASE;`nQUIT;`n" | Set-Content $probeSql -Encoding ascii
$probe = { param($Isql, $Sql, $Target, $User, $Password, $T0)
$started = Get-Date; $sw = [Diagnostics.Stopwatch]::StartNew()
$out = & $Isql -q -user $User -password $Password -i $Sql $Target 2>&1
'{0,8:n1}s probe SMALL {1,7} ms exit {2} {3}' -f ($started - $T0).TotalSeconds, $sw.ElapsedMilliseconds, $LASTEXITCODE, ((($out | ForEach-Object { "$_" }) -join ' ') -replace '\s+', ' ').Trim()
}
Write-Host ''
Write-Host ("T0: exclusive validation of BIG ({0} MB) starts; {1} waiter(s) at +{2}s then every {3}s; probe of SMALL every {4}s" -f [int]((Get-Item $big).Length / 1MB), $Waiters, $WaiterDelaySeconds, $WaiterStaggerSeconds, $ProbeIntervalSeconds)
$T0 = Get-Date
$val = Start-Process -FilePath $svc -ArgumentList @("${tcp}service_mgr", '-user', $User, '-password', $Password, '-action_repair', '-dbname', $big, '-rpr_validate_db', '-rpr_full') -NoNewWindow -PassThru
$waiterProcs = @(); $nextWaiter = $T0.AddSeconds($WaiterDelaySeconds); $left = $Waiters; $jobs = @()
while (-not $val.HasExited) {
$now = Get-Date
if ($left -gt 0 -and $now -ge $nextWaiter) {
$waiterProcs += Start-Process -FilePath $isql -ArgumentList @('-q', '-user', $User, '-password', $Password, '-i', $probeSql, "$tcp$big") -NoNewWindow -PassThru -RedirectStandardOutput (Join-Path $WorkDir "waiter-$($Waiters - $left + 1).out")
Write-Host ('{0,8:n1}s waiter {1} attaches to BIG' -f ($now - $T0).TotalSeconds, ($Waiters - $left + 1))
$left--; $nextWaiter = $now.AddSeconds($WaiterStaggerSeconds)
}
$jobs += Start-ThreadJob -ScriptBlock $probe -ArgumentList $isql, $probeSql, "$tcp$small", $User, $Password, $T0 -ThrottleLimit 128
foreach ($j in @($jobs | Where-Object State -eq 'Completed')) { Receive-Job $j | Write-Host; Remove-Job $j; $jobs = @($jobs | Where-Object Id -ne $j.Id) }
Start-Sleep -Seconds $ProbeIntervalSeconds
}
Write-Host ('{0,8:n1}s validation finished (exit {1})' -f ((Get-Date) - $T0).TotalSeconds, $val.ExitCode)
if ($left -gt 0) { Write-Warning "the validation ended before all $Waiters waiters had attached: nothing was tested. Raise -BigMB (the validation must last well past +$($WaiterDelaySeconds + ($Waiters - 1) * $WaiterStaggerSeconds) s)." }
$deadline = (Get-Date).AddSeconds(200)
while ($jobs -and (Get-Date) -lt $deadline) { foreach ($j in @($jobs | Where-Object State -ne 'Running')) { Receive-Job $j | Write-Host; Remove-Job $j; $jobs = @($jobs | Where-Object Id -ne $j.Id) }; Start-Sleep -Seconds 1 }
$waiterProcs | Where-Object { -not $_.HasExited } | ForEach-Object { Stop-Process -Id $_.Id -Force }
While a database is being validated exclusively (Services API
isc_action_svc_repairwithisc_spb_rpr_validate_db | isc_spb_rpr_full, i.e.gfix -v -full), attachments to that database wait for the validation to finish, which is expected. But as soon as a second attachment is waiting for the validated database, every new attachment to every other database on the same SuperServer hangs too, until the validation ends. One waiting attachment is harmless; the second one takes the whole server down for new connections. Existing attachments keep working. Clients see the attach hang for the duration of the validation (orconnection rejected by remote interfaceonceConnectionTimeoutexpires — 180 s by default).Reproduced on Windows x64 with the official zip kits (SuperServer, stock
firebird.confplus the port), TCP and XNET alike, on every version I tried:Version matrix
Two databases on the same server: BIG (12 GB, or 4 GB where noted) and SMALL (25 MB). Exclusive full validation of BIG starts at T0; K "waiters" attach to BIG (one
isqlprocess each, staggered 6 s); a probe attaches to SMALL every 2 s and runsSELECT 1 FROM RDB$DATABASE. Normal probe: 150-250 ms (process start included)."Blocked from 2nd waiter": in every run the first hung probe is the first one started after the second waiter's
isqlwas launched (within 50 ms), and every probe started until the validation's end hangs; they are all released within a second of the validation finishing. Probes started while only one waiter existed (up to 8 s of them) were normal.Variants on 5.0.4 (4 GB BIG), all with 2 waiters:
ConnectionTimeout = 20Reproduction
repro-attach-freeze.ps1(at the end of this issue) does all of this against any running server: creates BIG and SMALL, starts the validation withfbsvcmgr, starts the waiters, prints one line per probe. By hand it is five commands:What the threads are doing (full dump of firebird.exe 5.0.4.1812 during the hang, with the release PDBs)
Four kinds of threads, forming one chain:
The validator (service thread):
Validation::run → walk_database → walk_relation → walk_pointer_page → walk_data_page → CCH_fetch_page → PIO_read, insideJProvider::internalAttach(theisc_dpb_verifyattach).Waiter 1 (worker thread,
attach_database → JProvider::internalAttach → PAG_header → CCH_FETCH → CCH_fetch_lock → get_buffer → BufferDesc::addRef → SyncObject::lock → wait): blocked on the header page latch. Offline validation keeps the header page latched for its whole run —Validation::walk_databasefetches it withfetch_page(true, HEADER_PAGE, …)and releases it onlyif (vdr_flags & VDR_online)(validation.cpp 1617-1621 in v5.0.4).PAG_headeris called frominternalAttachbeforeinitGuard.leave()("Basic DBB initialization complete", jrd.cpp ~1899), so this thread is blocked while holding the database'sdbb_init_finimutex.Waiter 2 (worker thread,
attach_database → JProvider::internalAttach → initAttachment → RefMutexUnlock::enter → Database::ExistenceRefMutex::enter → Mutex::enter): blocked on thatdbb_init_fini.initAttachmentruns with the process-globaldbInitMutexheld (guardDbInit.enter()at jrd.cpp 1696 on Windows / 1742 elsewhere, released at 1760 only afterinitAttachmentreturns for an existing database), so this thread is blocked while holdingdbInitMutex.Every other attachment (5 worker threads in the dump, all probes of SMALL:
attach_database → JProvider::internalAttach → EnsureUnlock<Mutex>::enter → Mutex::enter): blocked ondbInitMutexatguardDbInit.enter().So: header page latch (validator) →
dbb_init_finiof BIG (waiter 1) →dbInitMutex(waiter 2) → everyone. With a single waiter the chain stops at the per-database mutex and nothing else is affected, which is why one waiter is harmless and two are not. The same code (theguardDbInit/initGuardordering aroundPAG_header, and the header page held by offline validation) is present in 3.0.14 and 4.0.7, matching the matrix.Killing the waiting clients does not help: the worker threads are inside the engine and never look at the socket; the server keeps their attachments (and the chain) until the validation ends.
What does and does not change it
firebird.confvsDefaultDbCachePages = 128K+FileSystemCacheThreshold = 100MConnectionTimeout= 20 (server and/or client), probes every 2 sConnectionTimeout= 20, probes every 30 sSQLSTATE = 08004 connection rejected by remote interface; the server keeps its socket in CLOSE_WAIT and its worker thread blocked until the validation endsTwo ways the client sees the same freeze, depending on where its connection blocks:
SecDbCache.cpp: released by a 10 s timer), authentication completes, the client getsop_accept, sendsop_attach, and its worker blocks ondbInitMutexinside the target database's attach. No timeout applies afterop_accept: the client hangs until the validation ends.JProvider::internalAttachand blocks ondbInitMutexbeforeop_acceptis sent. The client is still in the connect phase, itsConnectionTimeout(default 180 s) expires, and it reportsconnection rejected by remote interface. The server's worker stays blocked, so every such client leaves one blocked thread and one CLOSE_WAIT socket behind until the validation ends. With sparse connections this is what one sees: every client failing after 180 s while blocked threads pile up on the server.Environment
firebird -a), ServerMode = Super, one instance per version on its own port.repro-attach-freeze.ps1
PowerShell 7 script (creates the two databases, runs the validation, the waiters and the probe against any running server)