Skip to content

[#1149] Replay a rolled back PDB transaction until its conflict clears, bounded only by an optional db-txn-retry-time-limit - #1152

Open
vharseko wants to merge 1 commit into
OpenIdentityPlatform:masterfrom
vharseko:issue-1149-pdb-retry-window
Open

vharseko wants to merge 1 commit into
OpenIdentityPlatform:masterfrom
vharseko:issue-1149-pdb-retry-window

Conversation

@vharseko

@vharseko vharseko commented Oct 1, 2026

Copy link
Copy Markdown
Member

Fixes #1149

Problem

Since #937, PDBStorage.write() gives up after 10 attempts with result 80 (other). Ordinary concurrent ADD and
DELETE on a single suffix spend that cap: the JE vs PDB benchmark shows 6–8 such failures on PDB in every run, and
none on JE.

The analysis is in the issue.
In short:

  • The transactions collide on index keys shared by all entries, such as objectClass and the substrings of a
    common mail tail, for as long as those keys hold fewer IDs than index-entry-limit. Write throughput
    quadruples exactly when the run reaches 4000 live entries, on both backends, and no failure happens after that.
  • On persistit a rollback is the normal way two writers of the same key are resolved. The writer waits for the
    other transaction, and is rolled back if that one committed. A healthy write can therefore lose many races in a
    row. JE waits for the lock instead (je.lock.timeout=0), so the same cap is unreachable there.
  • The count of 10 was copied from JDBCStorage (Seek the primary key in the SQL Server upsert and retry a transaction conflict #867), where it bounds deadlock replays against a database that
    other writers may share. It was never measured for PDB.

Change

#921's guarantee becomes opt-in. With the default, a configuration change under the exclusive lock waits for its
conflict to clear, the way a JE writer waits for a lock. Under that lock no user write reaches the suffix's trees,
so a conflict there can only come from short-lived shared structures. Each rollback also means another transaction
committed, so the system keeps making progress. An operator who needs a hard bound sets the property.

Tests

PDBStorageTest:

Test Pins
testWriteOutlastsAnyNumberOfConflictsByDefault (new) 11 fast rollbacks with the default configuration, and the write commits. Master gives up on attempt 10
testRetryTimeLimitChangedWhileOpenAppliesToTheNextWrite (new) a limit set on an open storage ends the next write, with no restart asked for
testWriteIsReplayedOnceWhenTheFirstAttemptOutlastsTheRetryTimeLimit the attempt > 1 exemption
testWriteGivesUpOnTheRetryTimeLimitWhenAttemptsAreSlow the give-up happens on attempt 2; the conflict is suppressed; what the attempts wrote is rolled back
testExhaustedWriteNamesTheAttemptsItSpent the message names the attempts, the property and its value; getCause() is null

The tests whose conflict never clears call failIfReplayedPastTheLimit. Without it, a write that ignores the limit
would replay for as long as the default allows, and hang the test instead of failing it. The test of the attempt
cap and the 4-argument test constructor are gone, since the limit now comes from the configuration mock.

Mutants of PDBStorage, each run against the whole class:

Mutant Failing tests
attempt cap of 10 restored (master) testWriteOutlastsAnyNumberOfConflictsByDefault
a limit of 0 read as a window already spent 2
attempt > 1 exemption removed 4
limit never consulted 3
limit read once, from the configuration at construction testRetryTimeLimitChangedWhileOpenAppliesToTheNextWrite

Locally, -Pprecommit verify passes 158 tests with 0 failures: PDBStorageTest 38, PDBTestCase 39,
EncryptedPDBTestCase 39, PDBIndexConfidentialityChangeTest 5, ConfigChangeGivesUpTest 10,
ReplayedConfigChangeTest 24 and ReplayedOpenTest 3.

After the merge, the JE vs PDB benchmark should show no result-80 failures on PDB. The knee at 4000 live entries
stays, since it comes from how index keys are stored. The error lines the benchmark step log hides are #1150.

…l its conflict clears, bounded only by an optional db-txn-retry-time-limit

The attempt cap of 10 that OpenIdentityPlatform#937 took over from JDBCStorage failed ordinary concurrent ADD and DELETE with
result 80. On persistit a rollback is how two writers of the same key are resolved, and every entry rewrites
the index keys it shares with other entries while they stay below the index entry limit. The replays are now
bounded by time only, through a new pdb-backend property whose default of 0 replays without limit, the way a
JE writer waits for a lock, and the Storage.write() contract says so.

Fixes OpenIdentityPlatform#1149
@vharseko
vharseko requested a review from maximthomas October 1, 2026 14:17
@vharseko vharseko added bug concurrency Thread-safety / race-condition bugs java Changes to Java sources tests Test suites: fixing, enabling, un-disabling docs labels Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug concurrency Thread-safety / race-condition bugs docs java Changes to Java sources tests Test suites: fixing, enabling, un-disabling

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBStorage.write() fails ordinary concurrent ADD/DELETE with result 80: the attempt cap from #937 is spent under single-suffix write load

1 participant