[#1149] Replay a rolled back PDB transaction until its conflict clears, bounded only by an optional db-txn-retry-time-limit - #1152
Open
vharseko wants to merge 1 commit into
Conversation
…l its conflict clears, bounded only by an optional db-txn-retry-time-limit The attempt cap of 10 that OpenIdentityPlatform#937 took over from JDBCStorage failed ordinary concurrent ADD and DELETE with result 80. On persistit a rollback is how two writers of the same key are resolved, and every entry rewrites the index keys it shares with other entries while they stay below the index entry limit. The replays are now bounded by time only, through a new pdb-backend property whose default of 0 replays without limit, the way a JE writer waits for a lock, and the Storage.write() contract says so. Fixes OpenIdentityPlatform#1149
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1149
Problem
Since #937,
PDBStorage.write()gives up after 10 attempts with result 80 (other). Ordinary concurrent ADD andDELETE on a single suffix spend that cap: the
JE vs PDBbenchmark shows 6–8 such failures on PDB in every run, andnone on JE.
The analysis is in the issue.
In short:
objectClassand the substrings of acommon
mailtail, for as long as those keys hold fewer IDs thanindex-entry-limit. Write throughputquadruples exactly when the run reaches 4000 live entries, on both backends, and no failure happens after that.
other transaction, and is rolled back if that one committed. A healthy write can therefore lose many races in a
row. JE waits for the lock instead (
je.lock.timeout=0), so the same cap is unreachable there.JDBCStorage(Seek the primary key in the SQL Server upsert and retry a transaction conflict #867), where it bounds deadlock replays against a database thatother writers may share. It was never measured for PDB.
Change
PDBStorage.write(): the attempt cap (MAX_RETRIES) and the hard-coded 10 s window are removed. Only anew property ends the replays:
db-txn-retry-time-limit.je.lock.timeout=0. That is the behaviour before [#921] Bound the transaction replay of PDBStorage.write() #937.attempt > 1exemption is kept: an attempt that outlasts the limit is still replayed once.report what an operator would raise.
PDBBackendConfiguration.xmland02-config.ldif: the new advanced duration property (ms, lower limit0) and its attribute
ds-cfg-db-txn-retry-time-limit, OID1.3.6.1.4.1.60142.2.1.1.2, the next one after [#901] Make the replay retry budget of a replication domain configurable #944.Storage.write()contract: [#921] Bound the transaction replay of PDBStorage.write() #937 wrote "A replay must be bounded". It now says a replay may be bounded, ormay go on for as long as the conflict lasts, and that an engine resolving every conflict by a rollback should
bound it by time only.
db-checkpointer-wakeup-interval.#921's guarantee becomes opt-in. With the default, a configuration change under the exclusive lock waits for its
conflict to clear, the way a JE writer waits for a lock. Under that lock no user write reaches the suffix's trees,
so a conflict there can only come from short-lived shared structures. Each rollback also means another transaction
committed, so the system keeps making progress. An operator who needs a hard bound sets the property.
Tests
PDBStorageTest:testWriteOutlastsAnyNumberOfConflictsByDefault(new)testRetryTimeLimitChangedWhileOpenAppliesToTheNextWrite(new)testWriteIsReplayedOnceWhenTheFirstAttemptOutlastsTheRetryTimeLimitattempt > 1exemptiontestWriteGivesUpOnTheRetryTimeLimitWhenAttemptsAreSlowtestExhaustedWriteNamesTheAttemptsItSpentgetCause()is nullThe tests whose conflict never clears call
failIfReplayedPastTheLimit. Without it, a write that ignores the limitwould replay for as long as the default allows, and hang the test instead of failing it. The test of the attempt
cap and the 4-argument test constructor are gone, since the limit now comes from the configuration mock.
Mutants of
PDBStorage, each run against the whole class:testWriteOutlastsAnyNumberOfConflictsByDefaultattempt > 1exemption removedtestRetryTimeLimitChangedWhileOpenAppliesToTheNextWriteLocally,
-Pprecommit verifypasses 158 tests with 0 failures:PDBStorageTest38,PDBTestCase39,EncryptedPDBTestCase39,PDBIndexConfidentialityChangeTest5,ConfigChangeGivesUpTest10,ReplayedConfigChangeTest24 andReplayedOpenTest3.After the merge, the
JE vs PDBbenchmark should show no result-80 failures on PDB. The knee at 4000 live entriesstays, since it comes from how index keys are stored. The error lines the benchmark step log hides are #1150.