You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Since #937 (fix for #921), PDBStorage.WriteableStorageImpl.write() gives up after MAX_RETRIES = 10 attempts
and the operation fails with result code 80 (other):
pdb: backend 'userRoot' did not apply the transaction after 10 attempts in 2830 ms, the attempt cap being spent;
the last conflict was com.persistit.exception.RollbackException
This bound was designed for the configuration change paths, which hold an entry container's exclusive lock
across write(). But it applies to every write, and ordinary concurrent ADD/DELETE traffic in a single suffix
reaches it. Before #937 the same operations were replayed until they went through and succeeded.
Evidence: the JE vs PDB benchmark in build-docker
The benchmark runs 200 threads with rampup=0 for 150 s. Each iteration does ADD, SEARCH, COMPARE, MODIFY,
BIND and DELETE of its own mail=u_<thread>_<iter> entry under ou=People, then a READD of a UUID entry.
The figures below come from the benchmark-pdb-vs-je artifacts, parsed as CSV:
Every result-80 failure falls in the first ~20 s of the PDB run, between +0.5 s and +19.5 s.
Each failed operation took 1.8–3.9 s, which is the backoff ladder of 10 attempts (at most 5.5 s, ~2.8 s
expected). The bound that ran out is the attempt cap, not the 10 s window.
The cascade follows from each failed ADD: the entry is missing, so the COMPARE and MODIFY of that iteration
fail with 32, and its BIND fails with 49.
The master runs of build.yml show when this started. The benchmark has printed its error lines since #655
(06.07):
34603662889 (11.09) and 34692235136 (12.09), the first
runs after the merge: PDB prints 2 BIND | 49 | Invalid Credentials, the visible end of the same cascade.
The step log shows only the BIND line because of a separate defect in the benchmark script. .github/benchmark/compare-opendj.sh splits the JTL with awk -F','. The result-80 message and the DNs in the
32 messages contain commas, which shift the success column, so those rows never print. The err columns of
the step summary are read from statistics.json, and they show all of the errors.
Why JE is not affected by the same cap
JEStorage has the same MAX_RETRIES = 10 and the same 10 s window, added by #1065 (#1064). There the cap is
practically unreachable. je.lock.timeout is 0, so a writer waits for a lock instead of failing, and the only
conflict JE raises is a deadlock.
persistit works optimistically instead. Two transactions that write the same key collide, and one of them is
rolled back. The code already acknowledges this: the comment above the exhaustion WARN in PDBStorage.write()
says "a conflict is routine on the ordinary add and modify path of this engine". So on PDB, 10 attempts means
losing 10 races in a row for a hot key. When 200 writers start at once under one parent, that happens
reproducibly.
The sizing discussion in #937 went like this. The conflicts that drive the loop "come from other base DNs of
the same backend", "the page-level conflicts persistit reports resolve in microseconds", and 10 attempts
within 10 s "leaves a genuinely contended write ample room". The window was copied from JDBCStorage and was
not measured, because #921's question was about the configuration change paths. The data above is the
measurement #921 asked for. Under ordinary single-suffix load, the attempt cap runs out well before the window
does.
What it needs
Settle on a bound that a contended but healthy write does not spend, while the configuration change paths
still get a finite wait. Options, none of them measured yet:
Make the bounds configurable on pdb-backend, keeping the current defaults.
Find which key the ADD/DELETE transactions collide on, for example the parent's children count or an index
key shared by all entries, and reduce the contention there.
Whatever is chosen, the JE vs PDB benchmark should show zero result-80 failures on PDB, as it did before #937. PDBStorageTest.testWriteGivesUpAfterTheAttemptCap pins the shipped MAX_RETRIES and would need to follow.
Describe the bug
Since #937 (fix for #921),
PDBStorage.WriteableStorageImpl.write()gives up afterMAX_RETRIES = 10attemptsand the operation fails with result code 80 (
other):This bound was designed for the configuration change paths, which hold an entry container's exclusive lock
across
write(). But it applies to every write, and ordinary concurrent ADD/DELETE traffic in a single suffixreaches it. Before #937 the same operations were replayed until they went through and succeeded.
Evidence: the
JE vs PDBbenchmark inbuild-dockerThe benchmark runs 200 threads with
rampup=0for 150 s. Each iteration does ADD, SEARCH, COMPARE, MODIFY,BIND and DELETE of its own
mail=u_<thread>_<iter>entry underou=People, then a READD of a UUID entry.The figures below come from the
benchmark-pdb-vs-jeartifacts, parsed as CSV:expected). The bound that ran out is the attempt cap, not the 10 s window.
fail with 32, and its BIND fails with 49.
The master runs of
build.ymlshow when this started. The benchmark has printed its error lines since #655(06.07):
34692235136 (12.09), the first
runs after the merge: PDB prints
2 BIND | 49 | Invalid Credentials, the visible end of the same cascade.The step log shows only the BIND line because of a separate defect in the benchmark script.
.github/benchmark/compare-opendj.shsplits the JTL withawk -F','. The result-80 message and the DNs in the32 messages contain commas, which shift the
successcolumn, so those rows never print. Theerrcolumns ofthe step summary are read from
statistics.json, and they show all of the errors.Why JE is not affected by the same cap
JEStoragehas the sameMAX_RETRIES = 10and the same 10 s window, added by #1065 (#1064). There the cap ispractically unreachable.
je.lock.timeoutis 0, so a writer waits for a lock instead of failing, and the onlyconflict JE raises is a deadlock.
persistit works optimistically instead. Two transactions that write the same key collide, and one of them is
rolled back. The code already acknowledges this: the comment above the exhaustion WARN in
PDBStorage.write()says "a conflict is routine on the ordinary add and modify path of this engine". So on PDB, 10 attempts means
losing 10 races in a row for a hot key. When 200 writers start at once under one parent, that happens
reproducibly.
What #921 / #937 assumed
The sizing discussion in #937 went like this. The conflicts that drive the loop "come from other base DNs of
the same backend", "the page-level conflicts persistit reports resolve in microseconds", and 10 attempts
within 10 s "leaves a genuinely contended write ample room". The window was copied from
JDBCStorageand wasnot measured, because #921's question was about the configuration change paths. The data above is the
measurement #921 asked for. Under ordinary single-suffix load, the attempt cap runs out well before the window
does.
What it needs
Settle on a bound that a contended but healthy write does not spend, while the configuration change paths
still get a finite wait. Options, none of them measured yet:
the backoff ladder can spend in 10 s. PDBStorage.write() replays a rolled-back transaction without any bound, and configuration changes hold an entry container's exclusive lock across it #921 asked for a finite stall under the exclusive lock, and the window
already provides that.
pdb-backend, keeping the current defaults.key shared by all entries, and reduce the contention there.
Whatever is chosen, the
JE vs PDBbenchmark should show zero result-80 failures on PDB, as it did before #937.PDBStorageTest.testWriteGivesUpAfterTheAttemptCappins the shippedMAX_RETRIESand would need to follow.Related
JEStorage, unreachable there with the shipped configurationBIND"LDAP connection has been closed",ephemeral port exhaustion); it is unrelated to the backend