Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -248,6 +248,8 @@ Import task 20120917100628767 scheduled to start Sep 17, 2012 10:06:28 AM CEST
----
If write traffic to your directory service occurs in short bursts, and you use database backends of type `pdb`, you can potentially improve short-term performance during the bursts by increasing the `db-checkpointer-wakeup-interval` setting. This setting specifies the maximum length of time between attempts to write a checkpoint to the journal. Longer intervals allow more updates to accumulate in buffers before they are required to be written to disk. The transaction log is still written to disk, but the modified pages are kept in memory longer before being written. Longer intervals potentially cause recovery from an abrupt termination to take more time.

Database backends of type `pdb` resolve two concurrent updates of the same record by rolling one of them back and running it again. This is routine under concurrent write load, for example while index keys shared by many entries, such as their object classes, hold fewer entries than the index entry limit. By default, an update is run again for as long as it keeps being rolled back. The `db-txn-retry-time-limit` setting bounds that time: once it is spent, the operation fails with result code 80 (`other`). Set it only when a bounded wait matters more than the update, for example while you change the indexes or base DNs of a busy backend, which hold their suffix exclusively while they write. The change takes effect without a restart.


[#perf-import]
==== LDIF Import Settings
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,42 @@
</ldap:attribute>
</adm:profile>
</adm:property>
<adm:property name="db-txn-retry-time-limit" advanced="true">
<adm:synopsis>
Specifies how long a write transaction which the database rolled back is
run again before the operation which issued it fails.
</adm:synopsis>
<adm:description>
Persistit resolves two transactions which write the same record by rolling
one of them back once the other commits, and the transaction rolled back is
then run again. Under concurrent writes this is routine rather than a
failure: every entry added or deleted rewrites the index keys it shares
with other entries, such as its object classes, for as long as those keys
hold fewer entries than the index entry limit. A value of 0 runs the
transaction again for as long as it keeps being rolled back, the way a
writer of a JE backend waits for a lock. A positive value bounds that time:
once it is spent, the operation fails with the result code "other". The
bound is checked between two runs of the transaction, so a single run can
outlast it, and a transaction rolled back is always run once more. A bound
matters to the configuration changes of an index or of a base DN, which
write while they hold their suffix exclusively, so that every operation on
that suffix waits for them. Changes to this property take effect with the
next write transaction.
</adm:description>
<adm:default-behavior>
<adm:defined>
<adm:value>0s</adm:value>
</adm:defined>
</adm:default-behavior>
<adm:syntax>
<adm:duration base-unit="ms" lower-limit="0" />
</adm:syntax>
<adm:profile name="ldap">
<ldap:attribute>
<ldap:name>ds-cfg-db-txn-retry-time-limit</ldap:name>
</ldap:attribute>
</adm:profile>
</adm:property>
<adm:property name="disk-low-threshold" advanced="true">
<adm:synopsis>
Low disk threshold to limit database updates
Expand Down
9 changes: 8 additions & 1 deletion opendj-server-legacy/resource/schema/02-config.ldif
Original file line number Diff line number Diff line change
Expand Up @@ -4112,6 +4112,12 @@ attributeTypes: ( 1.3.6.1.4.1.60142.2.1.1.1
SYNTAX 1.3.6.1.4.1.1466.115.121.1.15
SINGLE-VALUE
X-ORIGIN 'OpenDJ Directory Server' )
attributeTypes: ( 1.3.6.1.4.1.60142.2.1.1.2
NAME 'ds-cfg-db-txn-retry-time-limit'
EQUALITY caseIgnoreMatch
SYNTAX 1.3.6.1.4.1.1466.115.121.1.15
SINGLE-VALUE
X-ORIGIN 'OpenDJ Directory Server' )
objectClasses: ( 1.3.6.1.4.1.26027.1.2.1
NAME 'ds-cfg-access-control-handler'
SUP top
Expand Down Expand Up @@ -5970,7 +5976,8 @@ objectClasses: ( 1.3.6.1.4.1.36733.2.1.2.23
ds-cfg-db-txn-no-sync $
ds-cfg-disk-full-threshold $
ds-cfg-disk-low-threshold $
ds-cfg-db-checkpointer-wakeup-interval )
ds-cfg-db-checkpointer-wakeup-interval $
ds-cfg-db-txn-retry-time-limit )
X-ORIGIN 'OpenDJ Directory Server' )
objectClasses: ( 1.3.6.1.4.1.36733.2.1.2.24
NAME 'ds-cfg-backend-index'
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -104,27 +104,6 @@ public final class PDBStorage implements Storage, Backupable, ConfigurationChang
{
private static final int IMPORT_DB_CACHE_SIZE = 32 * MB;

/**
* Number of attempts a {@link WriteableStorageImpl#write} makes before it propagates the conflict to the caller.
* <p>
* It is a budget of attempts and not of time, so it is only ever reached by the conflicts that report quickly.
* PersistIt reports a write-write conflict only once it has waited on it, up to
* {@code SharedResource.DEFAULT_MAX_WAIT_TIME} - a minute, which this backend never lowers - so a conflict slower
* to report than {@link #MAX_RETRY_WINDOW_NANOS} spends the whole window inside its first attempt, is granted the
* single replay that window's exemption guarantees, and gives up on the window after two attempts rather than
* after this many.
*/
static final int MAX_RETRIES = 10;

/**
* Wall-clock budget the replays of a {@link WriteableStorageImpl#write} may spend, in nanoseconds. It is checked
* between attempts, so an attempt already running is never interrupted, and never before one replay has been
* made: the loop returns after at most this window plus two attempts. It bounds the conflicts that are slow to
* report, which {@link #MAX_RETRIES} alone does not - an operation whose own work takes seconds would otherwise
* multiply that wait by the attempt count.
*/
static final long MAX_RETRY_WINDOW_NANOS = 10L * 1000L * 1000L * 1000L; //10 s

/**
* Upper bound of the random delay before the second attempt, in milliseconds; it doubles with every attempt. This
* is the bound of the flat sleep this loop took before it was bounded, so the first replay is delayed exactly as
Expand All @@ -135,6 +114,10 @@ public final class PDBStorage implements Storage, Backupable, ConfigurationChang
/** Upper bound the doubled delay is capped at, in milliseconds. */
private static final double MAX_SLEEP_ON_RETRY_MS = 1000.0;

/** Name of the property which bounds the replays of a {@link WriteableStorageImpl#write}, for the reports of it. */
private static final String RETRY_TIME_LIMIT_PROPERTY =
PDBBackendCfgDefn.getInstance().getDBTxnRetryTimeLimitPropertyDefinition().getName();

private static final String VOLUME_NAME = "dj";
private static final String JOURNAL_NAME = VOLUME_NAME + "_journal";
/** The buffer / page size used by the PersistIt storage. */
Expand Down Expand Up @@ -671,7 +654,10 @@ public void write(WriteOperation operation) throws Exception
{
final Transaction txn = db.getTransaction();
final long startedAt = System.nanoTime();
final long giveUpAt = startedAt + retryWindowNanos;
//read once, so that a change of the property applies from the next write on rather than to one in flight;
//0 replays for as long as the conflict lasts
final long retryTimeLimitMs = config.getDBTxnRetryTimeLimit();
final long giveUpAt = startedAt + TimeUnit.MILLISECONDS.toNanos(retryTimeLimitMs);
for (int attempt = 1;; attempt++)
{
final RollbackException conflict;
Expand Down Expand Up @@ -705,24 +691,26 @@ public void write(WriteOperation operation) throws Exception
// decided and slept for outside the try statement: the sleep used to run before the finally ended the
// rolled back transaction, holding it open for the whole backoff and lengthening the window every other
// writer collides with
//Bounded by time alone, and not by a count of attempts: a rollback is how persistit resolves two
//transactions writing the same key - it waits for the other one to end and rolls this one back if that
//one committed - so under concurrent writes to a hot key, such as an index key shared by entries below
//the index entry limit, a healthy write loses several of these races in a row (#1149)
//System.nanoTime() - giveUpAt is the overflow safe form of the comparison, and attempt > 1 keeps the
//window from ending the loop before a single replay: persistit reports a write-write conflict only once
//limit from ending the loop before a single replay: persistit reports a write-write conflict only once
//it has waited on it, up to SharedResource.DEFAULT_MAX_WAIT_TIME - a minute, which this backend never
//lowers - so one attempt can outlast the window on its own, and it is the attempt after that one which
//lowers - so one attempt can outlast the limit on its own, and it is the attempt after that one which
//is likeliest to succeed, the transaction that blocked it having just finished
//one clock sample for both, so that the elapsed time reported is the one the give up was decided on
final long now = System.nanoTime();
final boolean capSpent = attempt >= maxRetries;
if (capSpent || (attempt > 1 && now - giveUpAt >= 0))
if (retryTimeLimitMs > 0 && attempt > 1 && now - giveUpAt >= 0)
{
final long elapsedMs = TimeUnit.NANOSECONDS.toMillis(now - startedAt);
//which of the two bounds was spent, so that the config change paths - which report this as the trailing
//cause of a message of their own - say whether raising the attempts or the window is what would have helped
final String boundSpent = capSpent ? "attempt cap" : "retry window";
//names the property, so that the config change paths - which report this as the trailing cause of a
//message of their own - say what would have helped
final StorageRuntimeException spent = new StorageRuntimeException(
"pdb: backend '" + config.getBackendId() + "' did not apply the transaction after " + attempt
+ " attempts in " + elapsedMs + " ms, the " + boundSpent + " being spent; the last conflict was "
+ conflict);
+ " attempts in " + elapsedMs + " ms, the " + RETRY_TIME_LIMIT_PROPERTY + " of " + retryTimeLimitMs
+ " ms being spent; the last conflict was " + conflict);
// the conflict is suppressed rather than made the cause, because a cause is what every caller strips
// this message off with: write(WriteOperation) below unwraps a StorageRuntimeException that carries one
// and throws the cause in its place, and EntryContainer.throwAllowedExceptionTypes:1121 rethrows a
Expand All @@ -732,17 +720,17 @@ public void write(WriteOperation operation) throws Exception
spent.addSuppressed(conflict);
//warned once, at exhaustion only, unlike JDBCStorage which warns on every replay: a conflict is routine
//on the ordinary add and modify path of this engine and a line per replay would flood the log. It names
//the bound that was spent for the same reason the exception does, and it is the only rendering that can
//carry the stack of the conflict: stackTraceToSingleLineString, the form the config change paths report
//this exception with, walks the causes and never prints a suppressed exception
//the property that was spent for the same reason the exception does, and it is the only rendering that
//can carry the stack of the conflict: stackTraceToSingleLineString, the form the config change paths
//report this exception with, walks the causes and never prints a suppressed exception
logger.warn(LocalizableMessage.raw("pdb: giving up on the transaction of backend '%s' after %d attempts"
+ " in %d ms, the %s being spent: %s", config.getBackendId(), attempt, elapsedMs, boundSpent,
stackTraceToSingleLineString(conflict)));
+ " in %d ms, the %s of %d ms being spent: %s", config.getBackendId(), attempt, elapsedMs,
RETRY_TIME_LIMIT_PROPERTY, retryTimeLimitMs, stackTraceToSingleLineString(conflict)));
throw spent;
}
if (logger.isTraceEnabled())
{
logger.trace("pdb: replaying the transaction after %s, attempt %d of %d", conflict, attempt, maxRetries);
logger.trace("pdb: replaying the transaction after %s, attempt %d", conflict, attempt);
}
try
{
Expand Down Expand Up @@ -1030,10 +1018,6 @@ private StorageImpl newStorageImpl() {
private long configuredCacheSize;
private long reservedCacheSize;
private StorageStatus storageStatus = StorageStatus.working();
/** Attempt bound of a {@link WriteableStorageImpl#write}, {@link #MAX_RETRIES} outside the tests. */
private final int maxRetries;
/** Wall-clock bound of a {@link WriteableStorageImpl#write}, {@link #MAX_RETRY_WINDOW_NANOS} outside the tests. */
private final long retryWindowNanos;

/**
* Creates a new persistit storage with the provided configuration.
Expand All @@ -1046,35 +1030,8 @@ private StorageImpl newStorageImpl() {
*/
// FIXME: should be package private once importer is decoupled.
public PDBStorage(final PDBBackendCfg cfg, ServerContext serverContext) throws ConfigException
{
this(cfg, serverContext, MAX_RETRIES, MAX_RETRY_WINDOW_NANOS);
}

/**
* Creates a new persistit storage whose replay bounds are the given ones rather than {@link #MAX_RETRIES} and
* {@link #MAX_RETRY_WINDOW_NANOS}.
* <p>
* Only a test builds one of these, and it does so to stop the two bounds racing each other: with the shipped
* values a run of replays spends a random share of the window on backoff alone, so a test of the attempt cap
* can be ended by the window on a loaded machine, and a test of the window has to spend seconds of build time
* to reach it.
*
* @param cfg
* The configuration.
* @param serverContext
* This server instance context
* @param maxRetries
* Number of attempts a write makes before it propagates the conflict to the caller.
* @param retryWindowNanos
* Wall-clock budget the replays of a write may spend, in nanoseconds.
* @throws ConfigException if memory cannot be reserved
*/
PDBStorage(final PDBBackendCfg cfg, ServerContext serverContext, int maxRetries, long retryWindowNanos)
throws ConfigException
{
this.serverContext = serverContext;
this.maxRetries = maxRetries;
this.retryWindowNanos = retryWindowNanos;
backendDirectory = getBackendDirectory(cfg);
runningDirectoryPermissions = cfg.getDBDirectoryPermissions();
config = cfg;
Expand Down Expand Up @@ -1293,15 +1250,17 @@ public Importer startImport() throws ConfigException, StorageRuntimeException
/**
* {@inheritDoc}
* <p>
* A transaction the engine rolled back is replayed, bounded twice: by {@link #MAX_RETRIES} attempts and by the
* {@link #MAX_RETRY_WINDOW_NANOS} wall-clock window, whichever is spent first - except that the window alone
* never ends the replays before one has been made. It is bounded because the
* configuration change paths of the pluggable backend hold an entry container's exclusive lock across this
* method, and every reader of that suffix then waits - untimed and uninterruptibly - until it returns, so a
* conflict that never clears would park every worker thread of that suffix rather than fail one operation.
* A transaction the engine rolled back is replayed for as long as the {@code db-txn-retry-time-limit} of the
* backend allows - without limit when it is 0, its default - except that the limit never ends the replays before
* one has been made. A rollback is how persistit resolves two transactions writing the same key, so a healthy
* write under concurrent load can lose several of them in a row: a count of attempts failed such writes (#1149),
* and the replays are bounded by time alone. Without a limit a write waits for its conflict to clear the way a
* JE writer waits for a lock. A limit is for the configuration change paths of the pluggable backend, which hold
* an entry container's exclusive lock across this method: every reader of that suffix then waits - untimed and
* uninterruptibly - until it returns.
* <p>
* Once the bound is spent the conflict is reported as a {@link StorageRuntimeException} naming the backend, the
* attempts spent, the time they took and which of the two bounds ran out. It carries the conflict as a
* Once the limit is spent the conflict is reported as a {@link StorageRuntimeException} naming the backend, the
* attempts spent, the time they took and the property whose value ran out. It carries the conflict as a
* suppressed exception rather than as its cause: a cause is unwrapped below and thrown in its place, and
* {@code EntryContainer.throwAllowedExceptionTypes} likewise passes a {@link StorageRuntimeException} through
* untouched only while it has no cause. Given a cause, both hand the caller a bare RollbackException instead,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -75,10 +75,12 @@ public interface Storage extends Closeable
/**
* Executes a write operation. In case of a write operation rollback, implementations may replay the write
* operation rather than propagate the failure: a {@link WriteOperation} is required to be idempotent for
* exactly that reason. A replay must be bounded - by a number of attempts, by a window of time, or by both -
* so that a conflict which does not clear reaches the caller instead of being retried forever. The pluggable
* backend holds locks across this method, up to the exclusive lock of an entry container, and every thread
* waiting on one of those locks waits for as long as this method does.
* exactly that reason. A replay may be bounded - by a number of attempts, by a window of time, or by both - so
* that a conflict which does not clear reaches the caller, or may go on for as long as the conflict lasts, the
* way a writer of a lock based engine waits for a lock; an engine which resolves every conflict by a rollback
* should bound it by time only, since a healthy write under concurrent load loses several in a row. The
* pluggable backend holds locks across this method, up to the exclusive lock of an entry container, and every
* thread waiting on one of those locks waits for as long as this method does.
* <p>
* A caller that mutates state around this method must handle that bound being spent. Removing an entry from an
* in-memory map before the write so that a replay still finds the work to do, or reading configuration back out
Expand Down
Loading
Loading