Skip to content

PDBStorage.write() fails ordinary concurrent ADD/DELETE with result 80: the attempt cap from #937 is spent under single-suffix write load #1149

Description

@vharseko

Describe the bug

Since #937 (fix for #921), PDBStorage.WriteableStorageImpl.write() gives up after MAX_RETRIES = 10 attempts
and the operation fails with result code 80 (other):

pdb: backend 'userRoot' did not apply the transaction after 10 attempts in 2830 ms, the attempt cap being spent;
the last conflict was com.persistit.exception.RollbackException

This bound was designed for the configuration change paths, which hold an entry container's exclusive lock
across write(). But it applies to every write, and ordinary concurrent ADD/DELETE traffic in a single suffix
reaches it. Before #937 the same operations were replayed until they went through and succeeded.

Evidence: the JE vs PDB benchmark in build-docker

The benchmark runs 200 threads with rampup=0 for 150 s. Each iteration does ADD, SEARCH, COMPARE, MODIFY,
BIND and DELETE of its own mail=u_<thread>_<iter> entry under ou=People, then a READD of a UUID entry.
The figures below come from the benchmark-pdb-vs-je artifacts, parsed as CSV:

Run Side Samples Result 80, attempt cap spent Cascade
36845050319 (01.10, #1145) JE 649 030 0 —
PDB 715 578 6 (DELETE 3, ADD 2, READD 1) 2 × COMPARE 32, MODIFY 32, BIND 49
36738515054 (30.09, #1086) JE 538 600 0 —
PDB 581 568 8 (READD 4, DELETE 2, ADD 2) 2 × COMPARE 32, MODIFY 32, BIND 49
  • Every result-80 failure falls in the first ~20 s of the PDB run, between +0.5 s and +19.5 s.
  • Each failed operation took 1.8–3.9 s, which is the backoff ladder of 10 attempts (at most 5.5 s, ~2.8 s
    expected). The bound that ran out is the attempt cap, not the 10 s window.
  • The cascade follows from each failed ADD: the entry is missing, so the COMPARE and MODIFY of that iteration
    fail with 32, and its BIND fails with 49.

The master runs of build.yml show when this started. The benchmark has printed its error lines since #655
(06.07):

The step log shows only the BIND line because of a separate defect in the benchmark script.
.github/benchmark/compare-opendj.sh splits the JTL with awk -F','. The result-80 message and the DNs in the
32 messages contain commas, which shift the success column, so those rows never print. The err columns of
the step summary are read from statistics.json, and they show all of the errors.

Why JE is not affected by the same cap

JEStorage has the same MAX_RETRIES = 10 and the same 10 s window, added by #1065 (#1064). There the cap is
practically unreachable. je.lock.timeout is 0, so a writer waits for a lock instead of failing, and the only
conflict JE raises is a deadlock.

persistit works optimistically instead. Two transactions that write the same key collide, and one of them is
rolled back. The code already acknowledges this: the comment above the exhaustion WARN in PDBStorage.write()
says "a conflict is routine on the ordinary add and modify path of this engine". So on PDB, 10 attempts means
losing 10 races in a row for a hot key. When 200 writers start at once under one parent, that happens
reproducibly.

What #921 / #937 assumed

The sizing discussion in #937 went like this. The conflicts that drive the loop "come from other base DNs of
the same backend"
, "the page-level conflicts persistit reports resolve in microseconds", and 10 attempts
within 10 s "leaves a genuinely contended write ample room". The window was copied from JDBCStorage and was
not measured, because #921's question was about the configuration change paths. The data above is the
measurement #921 asked for. Under ordinary single-suffix load, the attempt cap runs out well before the window
does.

What it needs

Settle on a bound that a contended but healthy write does not spend, while the configuration change paths
still get a finite wait. Options, none of them measured yet:

  1. Let only the wall-clock window end the replay for PDB, and drop the attempt cap or raise it well beyond what
    the backoff ladder can spend in 10 s. PDBStorage.write() replays a rolled-back transaction without any bound, and configuration changes hold an entry container's exclusive lock across it #921 asked for a finite stall under the exclusive lock, and the window
    already provides that.
  2. Make the bounds configurable on pdb-backend, keeping the current defaults.
  3. Find which key the ADD/DELETE transactions collide on, for example the parent's children count or an index
    key shared by all entries, and reduce the contention there.

Whatever is chosen, the JE vs PDB benchmark should show zero result-80 failures on PDB, as it did before #937.
PDBStorageTest.testWriteGivesUpAfterTheAttemptCap pins the shipped MAX_RETRIES and would need to follow.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugconcurrencyThread-safety / race-condition bugsjavaChanges to Java sources

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions