Since 4.0 the Windows engine opens the database file with FILE_FLAG_OVERLAPPED
(src/jrd/os/win32/winnt.cpp, commit 3977345, "Optimization : use overlapped IO
for database files on Windows", 2020-03-31; 3.0 opens it with a plain synchronous
handle, g_dwExtraFlags = 0). On an overlapped handle the Windows cache manager
does no read-ahead, so every page the engine fetches becomes one synchronous 8 KB
disk read. Any sequential walk of a database that is not already in the OS file
cache — gfix -v -full, the online validation service, gbak -b, sweep — is then
bound by the storage's per-request latency instead of its bandwidth. Measured on the
same production server, same disk, same 62.7 GB database, two weeks apart:
373 KB per disk read at 386 MB/s on 3.0.12, 8.1 KB per disk read at 51 MB/s on
5.0.4; a full validation pass went from 135 s to 1,167 s (8.6×), and the sweep + gbak
remainder of the same step from 10–12 to 25–31 minutes (about 2.5×). Reproduced in a controlled lab on the same file with 3.0.12 and
5.0.4.1812, and isolated from Firebird entirely with a 40-line ReadFile test: the
same 8 KB sequential reads of a cold 4 GB region coalesce 15:1 through a synchronous
handle and 1:1 through an overlapped one.
The engine still opens the file this way on master (g_dwExtraFlags = FILE_FLAG_OVERLAPPED).
Version matrix
Production first: two SuperServer boxes (Windows Server 2022, Hyper-V virtual disks,
8 vCPU) migrated from 3.0.12 to 5.0.4.1812 in September; firebird.conf per our
package (DefaultDbCachePages = 128K, FileSystemCacheThreshold = 100M,
ServerMode = Super), page size 8192, header page buffers 131072 (1 GB cache).
The nightly gfix -v -full (Services API isc_action_svc_repair, rpr_validate_db | rpr_full) runs right after a 01:45 reboot, so the OS file cache is cold. Per-pass
means over the last 14 nights on each engine, from firebird.log
(Validation started / Validation finished):
| server |
database |
RAM |
3.0.12 |
5.0.4.1812 |
ratio |
| APP45 |
M, 62.7 GB |
56 GB |
135 s |
1,167 s |
8.6× |
| APP49 |
O, 62.5 GB |
40 GB |
96 s |
729 s |
7.6× |
| APP41 |
F, 22.1 GB |
40 GB |
25 s |
34 s |
1.4× |
What the disk saw during the pass (Prometheus, windows_exporter logical-disk
counters, 1-minute rate):
|
3.0.12 (APP45, 2026-09-10) |
5.0.4 (APP45, 2026-09-27) |
| bytes per disk read |
373 KB |
8.1 KB |
| disk reads/s |
2,400 |
6,500 |
| throughput |
386 MB/s |
51 MB/s |
sweep + gbak remainder of the step (firebird.log pass end → next database's pass start) |
10–12 min |
25–31 min (8 KB/read throughout) |
| the server's whole backup night, 35 databases (transcript start → end) |
24–37 min (19 nights) |
82–92 min (17 nights) |
The 22 GB database is not affected because it fits in the server's file cache. On the
two servers with a database larger than RAM, every database validated after the big
one that night is slow too (47–55 MB/s instead of 700–1,400 MB/s from cache), because
the big walk has evicted them.
Lab, same database file in both engines (the ODS 12 copy of O from the last
3.0 night and the ODS 13 file from two weeks later; official Windows x64 ZIP kits run
embedded; the fleet's three firebird.conf lines above; header buffers 131072; a
32 GB workstation, so the 61 GB file cannot be cached; Samsung 990 PRO NVMe; disk and
process counters sampled every 5 s):
| run |
engine |
operation |
time |
disk: reads/s, KB/read, MB/s |
engine page reads/s |
engine reads per disk read |
| R1 |
3.0.12 |
gfix -v -full, cold |
115 s |
10,400 · 78.9 · 525 |
68,800 |
6.6 |
| R6 |
3.0.12 |
same, repeat |
145 s |
11,800 · 55.8 · 540 |
57,600 |
4.9 |
| R2 |
5.0.4 |
gfix -v -full, cold |
612 s |
17,500 · 8.2 · 139 |
13,200 |
0.8 |
| R9 |
5.0.4 |
same, quiet disk |
378 s |
22,000 · 7.7 · 168 |
21,500 |
1.0 |
| R3 |
5.0.4 |
gfix -mend |
613 s |
23,900 · 11.0 · 225 |
13,800 |
0.6 |
| R4 |
5.0.4 |
online validation (fbsvcmgr action_validate) |
599 s |
24,000 · 8.1 · 188 |
13,700 |
0.6 |
| R10 |
5.0.4 |
-v -full, ParallelWorkers = MaxParallelWorkers = 4 |
373 s |
22,700 · 8.0 · 179 |
22,300 |
1.0 |
| R7 |
3.0.12 |
gbak -b -g |
483 s |
3,900 · 45.8 · 155 |
14,700 |
3.8 |
| R8 |
5.0.4 |
gbak -b -g (byte-identical 47.4 GB .gbk) |
696 s |
11,600 · 7.8 · 89 |
10,600 |
0.9 |
The ratio between engines is smaller on a local NVMe (2.6–5.3× depending on the concurrent load, ~0.046 ms per
request) than on the fleet's virtual disks (8–9× for the validation, ~0.15 ms per request): with 8 KB synchronous reads
the wall time is the number of pages times the per-request latency, whatever the
device's bandwidth. On 3.0 the OS read-ahead turned the same page walk into ~80–370 KB
requests. (A separate matrix we ran the same day for #9169 — 12 GB database, exclusive
-v -full — read 22 s on 3.0.14 and 100–120 s on 4.0.5, 4.0.7, 5.0.3 and 5.0.4, i.e.
the step is at 4.0, where the overlapped I/O arrived.)
Reproduction
-
Any database larger than the machine's RAM (or evict the file from the cache
between runs), page size 8192, on a Windows x64 kit. Run gfix -v -full on the same
file with a 3.0 kit (ODS 12 copy) and a 4.0/5.0 kit (ODS 13 copy) and watch
\LogicalDisk(_Total)\Avg. Disk Bytes/Read and Disk Read Bytes/sec: 3.0 reads in
50–400 KB requests at hundreds of MB/s; 4.0+ reads in 8 KB requests at whatever
1 / latency gives.
-
Without Firebird — the script at the end of this issue opens the same file twice
and reads 4 GB of never-touched regions in 8 KB sequential requests, once through a
plain handle and once through FILE_FLAG_OVERLAPPED + ReadFile +
GetOverlappedResult (what winnt.cpp does), reading the logical-disk raw
counters before and after each trial:
mode seconds MB/s disk reads KB/read coalesce
sync#1 5.1 805 34,717 121.0 15.1
overlapped#1 16.3 251 527,581 8.0 1.0
overlapped#2 16.6 247 529,590 8.0 1.0
sync#2 2.2 1,851 33,831 124.0 15.5
(coalesce = engine-style 8 KB requests per disk read.) The flag alone switches the
cache manager's read-ahead off; nothing else differs between the two trials.
Settings that do and do not help
| setting |
result |
ParallelWorkers = MaxParallelWorkers = 4 |
no effect on validation (R9 vs R10, 1.3 %; validation.cpp never reads att_parallel_workers, only sweep and index build do) |
FileSystemCacheThreshold = 100M (file cache on) vs the default 64K with a 128K page cache (file cache bypassed, FILE_FLAG_NO_BUFFERING) |
with the cache bypassed there is no read-ahead either, so it cannot be the way back |
DefaultDbCachePages / header page buffers 131072 (1 GB) |
the page cache does not matter for a single cold walk that touches every page once |
gfix -mend, online validation |
same page walk, same cost (R3, R4) |
| bigger page size |
not tried (would halve the request count, not fix the pattern) |
Where this was found
A fleet of 38 Windows SuperServer boxes migrating from 3.0.12 to 5.0.4 (17 done).
The nightly maintenance step of one 62 GB tenant (validate, sweep, gbak) went from
14–15 to 63–72 minutes and that server's backup night from ~31 to ~86 minutes,
and the 20-minute exclusive validation windows exposed #9169 nightly (two clients
waiting on the validated database freeze every attach on the server). We have moved
the night to the online validation service and dropped the second pass on our side,
but the gbak phase and any full validation stay 4–9× slower than on 3.0 for every
database that does not fit in RAM, and nothing in firebird.conf changes that.
#7508 (backup with -v 40–70 % slower on 4.0.2 than 2.5.9 on Windows, open since
2023) looks like the same mechanism seen on a smaller database.
Environment
- Production: Windows Server 2022 Standard (20348), Hyper-V VMs, 8 vCPU, 40–56 GB RAM,
SuperServer service; engines WI-V3.0.12.33787 (before) and WI-V5.0.4.1812 (after);
page size 8192; DefaultDbCachePages = 128K, FileSystemCacheThreshold = 100M,
ServerMode = Super; header page buffers 131072, sweep interval 0.
- Lab: Windows 11 Enterprise 22621, 24 logical CPUs, 32 GB RAM, Samsung 990 PRO NVMe;
official ZIP kits 3.0.12.33787 and 5.0.4.1812 run embedded (gfix, gbak,
fbsvcmgr loading the kit's engine in-process) with the same three config lines;
the same 61–62 GB production database in ODS 12.0 and ODS 13.1.
- 5.0.4.1812 is the latest release at the time of writing;
v5.0-release and master
carry the same g_dwExtraFlags = FILE_FLAG_OVERLAPPED.
Test-ReadAhead.ps1
PowerShell 7 script: 8 KB sequential reads of cold regions of any large file, synchronous vs overlapped handle, with the logical-disk raw counters before/after each trial
#Requires -Version 7
<#
.SYNOPSIS
Does Windows read-ahead survive FILE_FLAG_OVERLAPPED? Reads -SizeGb of a file in 8 KB
sequential requests, exactly like the Firebird page walk, through (a) a synchronous buffered
handle (Firebird 3's winnt.cpp) and (b) a FILE_FLAG_OVERLAPPED buffered handle with
ReadFile + GetOverlappedResult (Firebird 4/5's winnt.cpp). Each trial reads a DIFFERENT,
never-touched region of the file (cold), and the logical-disk raw counters before/after give
what the disk actually saw: reads, bytes, average bytes per read.
.EXAMPLE
.\Test-ReadAhead.ps1 -Path D:\big.fdb -SizeGb 4 -Trials 2 -StartGb 20
#>
param(
[Parameter(Mandatory)][string]$Path,
[double]$SizeGb = 4,
[int]$Trials = 2,
[double]$StartGb = 0
)
$ErrorActionPreference = 'Stop'
Add-Type -TypeDefinition @'
using System;
using System.Runtime.InteropServices;
public static class RawRead {
[StructLayout(LayoutKind.Sequential)] public struct OVERLAPPED { public IntPtr Internal; public IntPtr InternalHigh; public uint Offset; public uint OffsetHigh; public IntPtr hEvent; }
[DllImport("kernel32.dll", SetLastError=true, CharSet=CharSet.Unicode)] public static extern IntPtr CreateFile(string name, uint access, uint share, IntPtr sec, uint disp, uint flags, IntPtr tmpl);
[DllImport("kernel32.dll", SetLastError=true)] public static extern bool ReadFile(IntPtr h, byte[] buf, uint n, out uint read, ref OVERLAPPED ov);
[DllImport("kernel32.dll", SetLastError=true)] public static extern bool GetOverlappedResult(IntPtr h, ref OVERLAPPED ov, out uint read, bool wait);
[DllImport("kernel32.dll", SetLastError=true)] public static extern bool CloseHandle(IntPtr h);
[DllImport("kernel32.dll", SetLastError=true)] public static extern IntPtr CreateEvent(IntPtr sec, bool manual, bool init, string name);
public const uint GENERIC_READ = 0x80000000, FILE_SHARE_READ = 1, FILE_SHARE_WRITE = 2, OPEN_EXISTING = 3, FILE_ATTRIBUTE_NORMAL = 0x80, FILE_FLAG_OVERLAPPED = 0x40000000;
public const int ERROR_IO_PENDING = 997;
public static double Run(string path, bool overlapped, long start, long count, int block) {
uint flags = FILE_ATTRIBUTE_NORMAL | (overlapped ? FILE_FLAG_OVERLAPPED : 0);
IntPtr h = CreateFile(path, GENERIC_READ, FILE_SHARE_READ | FILE_SHARE_WRITE, IntPtr.Zero, OPEN_EXISTING, flags, IntPtr.Zero);
if (h == new IntPtr(-1)) throw new System.ComponentModel.Win32Exception(Marshal.GetLastWin32Error());
byte[] buf = new byte[block];
OVERLAPPED ov = new OVERLAPPED();
if (overlapped) ov.hEvent = CreateEvent(IntPtr.Zero, true, false, null);
var sw = System.Diagnostics.Stopwatch.StartNew();
for (long i = 0; i < count; i++) {
long off = start + i * block;
ov.Offset = (uint)(off & 0xFFFFFFFF); ov.OffsetHigh = (uint)(off >> 32); ov.Internal = IntPtr.Zero; ov.InternalHigh = IntPtr.Zero;
uint n;
bool ok = ReadFile(h, buf, (uint)block, out n, ref ov);
if (!ok) {
int err = Marshal.GetLastWin32Error();
if (overlapped && err == ERROR_IO_PENDING) ok = GetOverlappedResult(h, ref ov, out n, true);
if (!ok) throw new System.ComponentModel.Win32Exception(Marshal.GetLastWin32Error());
}
if (n != block) throw new Exception("short read at " + off);
}
sw.Stop();
if (overlapped) CloseHandle(ov.hEvent);
CloseHandle(h);
return sw.Elapsed.TotalMilliseconds;
}
}
'@
function Get-DiskRaw { $r = Get-CimInstance Win32_PerfRawData_PerfDisk_LogicalDisk -Filter "Name='_Total'"; [pscustomobject]@{ Reads = [double]$r.DiskReadsPersec; Bytes = [double]$r.DiskReadBytesPersec } }
$block = 8192
$count = [long]($SizeGb * 1GB / $block)
$offset = [long]($StartGb * 1GB)
'{0,-14} {1,8} {2,9} {3,10} {4,9} {5,9}' -f 'mode', 'seconds', 'MB/s', 'disk reads', 'KB/read', 'coalesce'
for ($t = 1; $t -le $Trials; $t++) {
foreach ($mode in @('sync', 'overlapped', 'overlapped', 'sync')[(($t - 1) * 2)..(($t - 1) * 2 + 1)]) {
$b0 = Get-DiskRaw
$ms = [RawRead]::Run($Path, ($mode -eq 'overlapped'), $offset, $count, $block)
$b1 = Get-DiskRaw
$reads = $b1.Reads - $b0.Reads; $bytes = $b1.Bytes - $b0.Bytes
'{0,-14} {1,8:N1} {2,9:N0} {3,10:N0} {4,9:N1} {5,9:N1}' -f "$mode#$t", ($ms/1000), ($SizeGb*1024/($ms/1000)), $reads, ($bytes/[math]::Max(1,$reads)/1KB), ($count/[math]::Max(1,$reads))
$offset += [long]($SizeGb * 1GB) # next trial reads a fresh region
}
}
Since 4.0 the Windows engine opens the database file with
FILE_FLAG_OVERLAPPED(
src/jrd/os/win32/winnt.cpp, commit 3977345, "Optimization : use overlapped IOfor database files on Windows", 2020-03-31; 3.0 opens it with a plain synchronous
handle,
g_dwExtraFlags = 0). On an overlapped handle the Windows cache managerdoes no read-ahead, so every page the engine fetches becomes one synchronous 8 KB
disk read. Any sequential walk of a database that is not already in the OS file
cache —
gfix -v -full, the online validation service,gbak -b, sweep — is thenbound by the storage's per-request latency instead of its bandwidth. Measured on the
same production server, same disk, same 62.7 GB database, two weeks apart:
373 KB per disk read at 386 MB/s on 3.0.12, 8.1 KB per disk read at 51 MB/s on
5.0.4; a full validation pass went from 135 s to 1,167 s (8.6×), and the sweep + gbak
remainder of the same step from 10–12 to 25–31 minutes (about 2.5×). Reproduced in a controlled lab on the same file with 3.0.12 and
5.0.4.1812, and isolated from Firebird entirely with a 40-line ReadFile test: the
same 8 KB sequential reads of a cold 4 GB region coalesce 15:1 through a synchronous
handle and 1:1 through an overlapped one.
The engine still opens the file this way on master (
g_dwExtraFlags = FILE_FLAG_OVERLAPPED).Version matrix
Production first: two SuperServer boxes (Windows Server 2022, Hyper-V virtual disks,
8 vCPU) migrated from 3.0.12 to 5.0.4.1812 in September;
firebird.confper ourpackage (
DefaultDbCachePages = 128K,FileSystemCacheThreshold = 100M,ServerMode = Super), page size 8192, header page buffers 131072 (1 GB cache).The nightly
gfix -v -full(Services APIisc_action_svc_repair,rpr_validate_db | rpr_full) runs right after a 01:45 reboot, so the OS file cache is cold. Per-passmeans over the last 14 nights on each engine, from
firebird.log(
Validation started/Validation finished):What the disk saw during the pass (Prometheus,
windows_exporterlogical-diskcounters, 1-minute rate):
firebird.logpass end → next database's pass start)The 22 GB database is not affected because it fits in the server's file cache. On the
two servers with a database larger than RAM, every database validated after the big
one that night is slow too (47–55 MB/s instead of 700–1,400 MB/s from cache), because
the big walk has evicted them.
Lab, same database file in both engines (the ODS 12 copy of O from the last
3.0 night and the ODS 13 file from two weeks later; official Windows x64 ZIP kits run
embedded; the fleet's three
firebird.conflines above; header buffers 131072; a32 GB workstation, so the 61 GB file cannot be cached; Samsung 990 PRO NVMe; disk and
process counters sampled every 5 s):
gfix -v -full, coldgfix -v -full, coldgfix -mendfbsvcmgr action_validate)-v -full,ParallelWorkers = MaxParallelWorkers = 4gbak -b -ggbak -b -g(byte-identical 47.4 GB .gbk)The ratio between engines is smaller on a local NVMe (2.6–5.3× depending on the concurrent load, ~0.046 ms per
request) than on the fleet's virtual disks (8–9× for the validation, ~0.15 ms per request): with 8 KB synchronous reads
the wall time is the number of pages times the per-request latency, whatever the
device's bandwidth. On 3.0 the OS read-ahead turned the same page walk into ~80–370 KB
requests. (A separate matrix we ran the same day for #9169 — 12 GB database, exclusive
-v -full— read 22 s on 3.0.14 and 100–120 s on 4.0.5, 4.0.7, 5.0.3 and 5.0.4, i.e.the step is at 4.0, where the overlapped I/O arrived.)
Reproduction
Any database larger than the machine's RAM (or evict the file from the cache
between runs), page size 8192, on a Windows x64 kit. Run
gfix -v -fullon the samefile with a 3.0 kit (ODS 12 copy) and a 4.0/5.0 kit (ODS 13 copy) and watch
\LogicalDisk(_Total)\Avg. Disk Bytes/ReadandDisk Read Bytes/sec: 3.0 reads in50–400 KB requests at hundreds of MB/s; 4.0+ reads in 8 KB requests at whatever
1 / latencygives.Without Firebird — the script at the end of this issue opens the same file twice
and reads 4 GB of never-touched regions in 8 KB sequential requests, once through a
plain handle and once through
FILE_FLAG_OVERLAPPED+ReadFile+GetOverlappedResult(whatwinnt.cppdoes), reading the logical-disk rawcounters before and after each trial:
(
coalesce= engine-style 8 KB requests per disk read.) The flag alone switches thecache manager's read-ahead off; nothing else differs between the two trials.
Settings that do and do not help
ParallelWorkers = MaxParallelWorkers = 4validation.cppnever readsatt_parallel_workers, only sweep and index build do)FileSystemCacheThreshold = 100M(file cache on) vs the default 64K with a 128K page cache (file cache bypassed,FILE_FLAG_NO_BUFFERING)DefaultDbCachePages/ header page buffers 131072 (1 GB)-mend, online validationWhere this was found
A fleet of 38 Windows SuperServer boxes migrating from 3.0.12 to 5.0.4 (17 done).
The nightly maintenance step of one 62 GB tenant (validate, sweep, gbak) went from
14–15 to 63–72 minutes and that server's backup night from ~31 to ~86 minutes,
and the 20-minute exclusive validation windows exposed #9169 nightly (two clients
waiting on the validated database freeze every attach on the server). We have moved
the night to the online validation service and dropped the second pass on our side,
but the gbak phase and any full validation stay 4–9× slower than on 3.0 for every
database that does not fit in RAM, and nothing in
firebird.confchanges that.#7508 (backup with
-v40–70 % slower on 4.0.2 than 2.5.9 on Windows, open since2023) looks like the same mechanism seen on a smaller database.
Environment
SuperServer service; engines WI-V3.0.12.33787 (before) and WI-V5.0.4.1812 (after);
page size 8192;
DefaultDbCachePages = 128K,FileSystemCacheThreshold = 100M,ServerMode = Super; header page buffers 131072, sweep interval 0.official ZIP kits 3.0.12.33787 and 5.0.4.1812 run embedded (
gfix,gbak,fbsvcmgrloading the kit's engine in-process) with the same three config lines;the same 61–62 GB production database in ODS 12.0 and ODS 13.1.
v5.0-releaseandmastercarry the same
g_dwExtraFlags = FILE_FLAG_OVERLAPPED.Test-ReadAhead.ps1
PowerShell 7 script: 8 KB sequential reads of cold regions of any large file, synchronous vs overlapped handle, with the logical-disk raw counters before/after each trial