[#1119] Serve each LDAP connection handler with a transport of its own, built from its configuration - #1125
Conversation
maximthomas
left a comment
There was a problem hiding this comment.
praise: The handler now applies its own transport settings, and they are checked on real sockets.
LDAPConnectionHandler2.newTransport()keeps what the SDK's shared server transport had (SameThreadIOStrategy, a pooled memory manager, one selector runner per selector thread). Only the configured values change, andMemoryManagerHolderkeeps the 3% heap pre-allocation from being multiplied per handler.LDAPConnectionHandler2TransportTestCasereadsSO_KEEPALIVE,TCP_NODELAY,SO_REUSEADDRand the write buffer from the server-side sockets of a real handler. It passes 12/12 at d7ddc56 in a local failsafe run, withopendj-grizzlyand the config XML built from the head.
issue (non-blocking): A TransportDrain waits until the whole handler has no connections, not until its own transport has none.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:166-171, :768-783, :830
connectionClosed() tests clientConnections.isEmpty(), and that set holds the connections of every transport of the instance, the live one included. One instance can run two transports. Take an LDAPS handler with clients and set ssl-cert-nickname to an alias the key store does not have. The change is accepted, because createSSLContext(config, false) only logs (:1164-1169). The apply then disables the instance through disableAndWarnIfUseSSL (:1122). ds-cfg-enabled stays true, so ConnectionHandlerConfigManager keeps the instance. run() drains T1, and after the alias is set back it starts T2 in the same instance. From then on, T1 and its numRequestHandlers selector threads stay alive until the handler has no connections at all, which on a busy handler never happens. Each such cycle adds one more transport. Clients are not affected. The class javadoc and the description both say the transport is shut down "after the last one closes". Traced, not run.
// startListener(): record which connections this transport accepts
transport = newTransport();
final Collection<ClientConnection> accepted = ConcurrentHashMap.newKeySet();
transportConnections = accepted;
// ... in apply(): clientConnections.add(conn); accepted.add(conn);
// ... in the three event methods: connectionClosed(conn, accepted);
private void connectionClosed(ClientConnection connection, Collection<ClientConnection> accepted) {
accepted.remove(connection);
clientConnections.remove(connection);
for (TransportDrain drain : drains) {
drain.connectionClosed();
}
}
// stopListener(): new TransportDrain(transport, transportConnections)
// TransportDrain keeps that collection as `accepted`, tests `accepted.isEmpty()` in connectionClosed(),
// and disconnects `accepted` in processServerShutdown().issue (non-blocking): Whether a handler's connections get the server-shutdown notice depends on a late isShuttingDown() read, and an in-core restart can race it.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:777, :439-441
stopListener() runs on the handler thread up to 1 s after finalizeConnectionHandler, because of StaticUtils.sleep(1000) in run(). DirectoryServer.shutDown never joins that thread. restart() then calls getNewInstance (DirectoryServer.java:1087), which creates a new instance with shuttingDown=false. If the late read sees that instance, the old handler's connections take drain.start() instead: no notice, they stay open across the restart, and the drain registers with the new server. If the read lands between :1087 and bootstrapServer (:1173), shutdownListeners is still null, so registerShutdownListener (:4103) throws an NPE on the handler thread. At BASE, whether connections survived a restart already depended on timing. What is new is that the promised notice can be skipped and that the NPE is possible. Not run: this needs a restart with a Handler2 connection open, repeated until the late read shows up.
/** Whether the server was shutting down when this handler was finalized. */
private volatile boolean serverShuttingDown;
@Override
public void finalizeConnectionHandler(LocalizableMessage finalizeReason) {
// DirectoryServer.shutDown sets the flag before it finalizes the handlers, on this thread.
serverShuttingDown = DirectoryServer.getInstance().isShuttingDown();
closeReason = finalizeReason;
// ...
}
// stopListener()
if (serverShuttingDown) {
drain.processServerShutdown(closeReason);
} else {
drain.start();
}issue (non-blocking): At server shutdown the drain calls shutdownNow() straight after the disconnects, so a notice queued behind a write backlog may be dropped.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:179-189
For an idle client, Grizzly writes the notice on the calling thread, and serverShutdownEndsTheConnectionsOfAStoppedHandler pins that case. Now take a client that has stopped reading a large result, with roughly 0.5 MiB or more unread on loopback. Its notice is queued behind the backlog, and shutdownNow() terminates the connection without waiting for the queue. That client gets EOF or a reset instead of the notice. The connection ends either way. Not traced into AbstractNIOAsyncQueueWriter or NIOTransport.finalizeShutdown at the Grizzly version the build resolves, and not reproduced. If a queued write can be dropped, wait a bounded time for the disconnect writes, or shut the transport down gracefully, before shutdownNow().
suggestion (non-blocking): No case checks that stopping a handler with no open connection shuts its transport down.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:268-281
This is the path taken whenever a handler with no clients is disabled, deleted or restarted: TransportDrain.start() ends with connectionClosed(), which shuts the transport down at once. Both cases that call assertThreadsStop stop their handler with a connection still open, so they end the drain some other way. The mutant that deletes that connectionClosed() call leaks the selector threads on every idle stop, and the class stays green.
private static void assertSelectorThreads(LDAPConnectionHandlerCfg config, int expected) throws Exception
{
final LDAPConnectionHandler2 handler = start(config);
final String threadPrefix = selectorThreadPrefix(handler);
try
{
final TCPNIOTransport transport = transport(handler);
assertEquals(transport.getSelectorRunnersCount(), expected, "selector runners");
assertEquals(transport.getKernelThreadPoolConfig().getMaxPoolSize(), expected, "selector threads");
}
finally
{
stop(handler);
}
assertThreadsStop(threadPrefix);
}suggestion (non-blocking): The server-shutdown test calls the drain directly, so neither the drain's registration nor stopListener's isShuttingDown() arm is pinned.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:226-229
Two mutants survive the whole class. The first deletes DirectoryServer.registerShutdownListener(this) from TransportDrain.start(). The second replaces the isShuttingDown() arm with drain.start(). The test's stop() never makes the server shut down, so stopListener() always takes drain.start(). No other suite asserts a shutdown-time notice on a Handler2 connection.
final Collection<?> drains = (Collection<?>) field(handler, "drains");
assertEquals(drains.size(), 1, "no transport is left serving the connection of the stopped handler");
final Object drain = drains.iterator().next();
assertTrue(((Collection<?>) field(DirectoryServer.getInstance(), "shutdownListeners")).contains(drain),
"the drain is not registered as a shutdown listener");
((ServerShutdownListener) drain).processServerShutdown(STOP_REASON);Pin for the arm: a case that holds a raw socket on the test server's own LDAP port (config.ldif runs LDAPConnectionHandler2 there), calls TestCaseUtils.restartServer(), and asserts the notice-of-disconnection OID arrives. The same case measures the restart race above.
suggestion (non-blocking): No case covers a failed listen, where the transport started before the bind fails.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:819-821, :944-945
startListener() starts the transport before new LDAPListener(...) binds. When the bind fails, the transport is drained only by the transport != null arm of stopListener(). The mutant that returns from stopListener() when listener == null leaks two transports per consecutive-failure cycle, and nothing fails. Both initializeConnectionHandler and isConfigurationAcceptable check the port first, so the test has to take the port between initialization and start.
@Test
public void failedListenLeavesNoSelectorThreads() throws Exception
{
final int port = TestCaseUtils.findFreePort();
final LDAPConnectionHandler2 handler = new LDAPConnectionHandler2();
handler.initializeConnectionHandler(DirectoryServer.getInstance().getServerContext(),
configuration(port, true, true, false, 2));
try (ServerSocket taken = new ServerSocket())
{
taken.bind(new InetSocketAddress("127.0.0.1", port));
handler.start();
final long deadline = System.currentTimeMillis() + TIMEOUT_MS;
while ((Boolean) field(handler, "enabled"))
{
assertTrue(System.currentTimeMillis() < deadline, "the handler did not give up listening");
Thread.sleep(50);
}
assertThreadsStop(handler.getConnectionHandlerName() + " Request Handler");
}
finally
{
stop(handler);
}
}suggestion (non-blocking): No case pins that num-request-handlers takes precedence over org.forgerock.opendj.transport.selectors.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:143-147
getNumRequestHandlers(config) is correct, but no case sets both the property and num-request-handlers. A mutant that lets the property win whenever it is set stays green.
@Test
public void numRequestHandlersTakesPrecedenceOverTheSelectorsProperty() throws Exception
{
final String saved = System.getProperty(SELECTORS_PROPERTY);
System.setProperty(SELECTORS_PROPERTY, "13");
try
{
assertSelectorThreads(configuration(TestCaseUtils.findFreePort(), true, true, true, 3), 3);
}
finally
{
restore(saved);
}
}suggestion (non-blocking): The notice assertion accepts any notice of disconnection, whatever its result code and message.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:239-240
LDAPClientConnection2.disconnect sends that OID for every DisconnectReason. So a drain that passes another reason, or a null message (which sends the generic text), stays green.
final String text = new String(received.toByteArray(), StandardCharsets.ISO_8859_1);
assertTrue(text.contains(NOTICE_OF_DISCONNECTION_OID), "the connection was closed without a notice of disconnection");
assertTrue(text.contains(STOP_REASON.toString()), "the notice does not carry the shutdown reason");Or: to also pin the result code (UNAVAILABLE, 52), decode the response with the SDK's LDAP reader.
suggestion (non-blocking): On 4-vCPU CI runners the processor-based selector count is 2, so the test cannot tell it from a constant 2.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:165-175
Math.max(2, cpus / 2) is 2 on any host with 5 or fewer CPUs, so a new-delegation mutant that passes 2 instead of null stays green on ubuntu-latest. The arm itself is BASE code. At most, add a comment saying the case is decisive only on hosts with 6 or more CPUs.
note (non-blocking): Change not described in the PR:
opendj-maven-plugin/src/main/resources/config/xml/org/forgerock/opendj/server/config/LDAPConnectionHandlerConfiguration.xml:91-108,:124-141,:396-406— the XML is shared, so the newcomponent-restartonuse-tcp-keep-alive,use-tcp-no-delayandbuffer-sizealso applies to handlers still on the legacyorg.opends.server.protocols.ldap.LDAPConnectionHandler, the XML's defaultjava-class. That class reads these values for each accepted connection, so dsconfig now reports a restart it does not need.
…onnections, and let queued notices out before shutting it down Review round 1 of OpenIdentityPlatform#1125: - A TransportDrain tests and disconnects only the connections its own transport accepted, not every connection of the handler, so a handler that stops and starts listening again in one instance no longer keeps the old transport alive for the new one's clients. - Whether the server is shutting down is recorded in finalizeConnectionHandler instead of being read when the handler thread stops the listener, up to a second later, when an in-core restart may already have made the next server instance current. A drain ended on that path never touches the shutdown listeners. - Before shutdownNow(), a transport waits up to 2 seconds for the Grizzly connections it accepted to close, tracked by a ConnectionProbe, so a notice of disconnection queued behind unread data is not dropped. - Tests: two transports in one handler, a notice queued behind unread data, an in-core restart of the test server, a failed listen, num-request-handlers over the system property, the idle stop in the selector cases, the drain's registration, and the decoded notice (OID, result code, message).
d7ddc56 to
70ec861
Compare
|
All ten points taken in 70ec861, on top of the first commit rebased onto master d0ce1d7. #1116 has been merged, and the rebase dropped only the XML header year change, which #1116 already made. 1. A drain waits for the whole handler. Fixed as suggested.
2. Late 3. 4. Idle stop. 5. Registration and the shutdown arm. 6. Failed listen. 7. Precedence. 8. Notice content. 9. Processor count. Comment added: the case is decisive only on hosts with 6 or more CPUs. 10. Checks. Directly with TestNG:
Two survive, both parts of the restart race that a test cannot hit on demand: going back to the late |
maximthomas
left a comment
There was a problem hiding this comment.
praise: Round 2 fixes all three round-1 findings where they occur, and each fix has a case that exercises it.
- Each
TransportDrainnow waits on and disconnects only its own transport'sacceptedset, so a stop/start in one instance no longer keeps the old transport alive (drainEndsWithTheConnectionsOfItsOwnTransport). serverShuttingDownis now recorded infinalizeConnectionHandlerinstead of being read late on the handler thread, and theOpenConnectionsprobe lets a notice queued behind unread data go out beforeshutdownNow().serverShutdownDeliversANoticeQueuedBehindUnreadDatafills the Grizzly write queue to prove it, and passed locally in 1.3 s.- 17 of the 18 cases in
LDAPConnectionHandler2TransportTestCasepassed locally at 70ec861.
issue (non-blocking): serverRestartEndsTheConnectionsOfAListeningHandler fails on macOS whenever the in-core restart takes longer than about 60 s.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:327-337, :520
The case reads the notice only after TestCaseUtils.restartServer() returns, so the notice and the server's FIN wait in the client's receive buffer for the whole restart. In a local macOS run the server sent the notice 1 s into the shutdown (DISCONNECT conn=10 reason="Server Shutdown"). The restart then took 95 s, most of it a reverse-DNS stall at start-up on that machine, and the read failed with java.net.SocketException: Connection reset (tests=18, failures=1). macOS resets a loopback connection whose server side has closed once net.inet.tcp.fin_timeout (60 s) expires, and the client's unread data is lost. A standalone probe confirms it: a read after 5 s and after 50 s returned the data, a read after 75 s got Connection reset. Linux cells are expected to be green, but this is not confirmed yet. Reading while the restart runs removes the dependency on how long it takes:
try (Socket client = new Socket("127.0.0.1", TestCaseUtils.getServerLdapPort()))
{
awaitServerConnection(client.getLocalPort());
// Read while the server restarts: macOS resets a closed loopback connection after
// net.inet.tcp.fin_timeout (60 s) and drops what the client has not read yet.
final ExecutorService reader = Executors.newSingleThreadExecutor();
try
{
final Future<?> notice = reader.submit(() -> {
assertNoticeOfDisconnection(client, INFO_CONNHANDLER_CLOSED_BY_SHUTDOWN.get());
return null;
});
TestCaseUtils.restartServer();
notice.get(TIMEOUT_MS, TimeUnit.MILLISECONDS);
}
finally
{
reader.shutdownNow();
}
}issue (non-blocking): If a handler is disabled or deleted less than a second before an in-core restart, its drain registers after shutDown has already notified the shutdown listeners.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:203-208, :512, :852
A disable or delete finalizes the handler with serverShuttingDown=false and deregisters it, so shutDown never finalizes it again. Its thread leaves sleep(1000) up to a second later and takes drain.start(). If shutDown's listener loop (DirectoryServer.java:4193-4195) has already run by then, the drain is never notified. Its clients get no notice, and the drained transport keeps serving them across the restart, which is also what happened at BASE. In an embedded server (daemon threads, so the swap does not wait for the handler thread), registerShutdownListener can also hit the null shutdownListeners between getNewInstance and bootstrap. The NPE then ends the thread before connectionClosed() runs. Re-checking after registering closes the first road:
private void start() {
drains.add(this);
registered = true;
DirectoryServer.registerShutdownListener(this);
if (DirectoryServer.getInstance().isShuttingDown()) {
// Registered after shutDown notified its listeners: end the connections as it would have.
processServerShutdown(closeReason);
} else {
connectionClosed();
}
}issue (non-blocking): TransportDrain.start() adds the drain to drains before it registers it, so if the last connection closes in between, a finished drain stays registered as a shutdown listener.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:203-208, :236-246
Suppose a selector thread closes the drained transport's last connection between drains.add(this) and registerShutdownListener(this). finish() then wins the CAS and either skips deregistration or removes a listener that is not registered yet. The handler thread then registers the finished drain, which stays in DirectoryServer.shutdownListeners with the handler and the shut-down transport until the next shutdown. With the fix below, whichever of the two runs last removes the listener:
DirectoryServer.registerShutdownListener(this);
if (finished.get()) {
// The last connection closed between drains.add and the registration, and finished the drain first.
DirectoryServer.deregisterShutdownListener(this);
}suggestion (non-blocking): No test covers the close half of OpenConnections. A drain that always waits the full 2 s still passes every case.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:162-169, test :313-314
If onCloseEvent does nothing, or awaitClosed becomes a flat 2 s wait, every drain and every server shutdown takes 2 s longer, and every case still passes. The queued-notice case joins the drain thread for TIMEOUT_MS (10 s), and no case bounds elapsed time below that.
assertNoticeOfDisconnection(client, STOP_REASON);
// The client has read everything and the server has closed: the drain must not wait out its 2 s bound.
shutdown.join(1000);
assertFalse(shutdown.isAlive(), "the drain is still waiting for a closed connection");Pin: with a no-op onCloseEvent the drain thread is still in awaitClosed at about 1.3 s, so the case fails. At HEAD the thread ends about 0.3 s after the notice is read.
suggestion (non-blocking): serverRestartEndsTheConnectionsOfAListeningHandler catches a revert of the finalize-time serverShuttingDown only when the handler thread wakes after getNewInstance.
opendj-server-legacy/src/test/java/org/opends/server/protocols/ldap/LDAPConnectionHandler2TransportTestCase.java:322-337
The embedded test server forces daemon threads, so getNewInstance does not wait for the handler thread. A revert to DirectoryServer.getInstance().isShuttingDown() in stopListener therefore takes the same road as HEAD whenever the thread wakes while the old instance is still shutting down. In the local run the notice went out at 16:09:15 and the old instance stopped at 16:09:18, so the thread read the flag about 3 s before the swap, and the revert would have passed. To pin the fix, the test needs a way to hold the handler thread until the swap has happened. Otherwise, the javadoc should drop "after the server has moved on", because the case does not force that order.
suggestion (non-blocking): No case reaches the catch arm in newTransport, which shuts down a transport whose start() failed.
opendj-server-legacy/src/main/java/org/forgerock/opendj/reactive/LDAPConnectionHandler2.java:888-893
In failedListenLeavesNoSelectorThreads the listen fails at the GrizzlyLDAPListener bind, after start() has succeeded, so the drain cleans up and the catch never runs.
Pin: there is no cheap one. It would need a way to make TCPNIOTransport.start() throw, and the arm is three lines and correct by reading, so it is fine to leave.
… transport of its own, built from its configuration LDAPConnectionHandler2 ran every listener on the transport the SDK shares across the JVM, so use-tcp-keep-alive, use-tcp-no-delay, buffer-size and num-request-handlers had no effect and allow-tcp-reuse-address reached only the pre-check. - GrizzlyLDAPListener takes a transport through the new GRIZZLY_TRANSPORT option and leaves it running when it is closed. - LDAPConnectionHandler2 starts a transport per listener: num-request-handlers selector threads (else org.forgerock.opendj.transport.selectors, else chosen from the processors), SO_REUSEADDR on the listen socket from allow-tcp-reuse-address, and buffer-size as the write buffer. SO_KEEPALIVE and TCP_NODELAY reach every accepted socket through the listener options. The transports share one pooled memory manager. - Stopping the listener leaves the connections it accepted open: its transport keeps serving them and is shut down after the last one closes. When the server shuts down, they are ended as a server shutdown, with a notice of disconnection, before the transport is shut down. - use-tcp-keep-alive, use-tcp-no-delay and buffer-size of the LDAP connection handler now require a component restart. Fixes OpenIdentityPlatform#1119
…onnections, and let queued notices out before shutting it down Review round 1 of OpenIdentityPlatform#1125: - A TransportDrain tests and disconnects only the connections its own transport accepted, not every connection of the handler, so a handler that stops and starts listening again in one instance no longer keeps the old transport alive for the new one's clients. - Whether the server is shutting down is recorded in finalizeConnectionHandler instead of being read when the handler thread stops the listener, up to a second later, when an in-core restart may already have made the next server instance current. A drain ended on that path never touches the shutdown listeners. - Before shutdownNow(), a transport waits up to 2 seconds for the Grizzly connections it accepted to close, tracked by a ConnectionProbe, so a notice of disconnection queued behind unread data is not dropped. - Tests: two transports in one handler, a notice queued behind unread data, an in-core restart of the test server, a failed listen, num-request-handlers over the system property, the idle stop in the selector cases, the drain's registration, and the decoded notice (OID, result code, message).
…ain starts, and bound how long a closed drain runs Review round 2 of OpenIdentityPlatform#1125: - A handler remembers the server instance it was initialized for, and a TransportDrain checks whether that instance is shutting down, instead of a flag recorded at finalization. A handler disabled or deleted just before an in-core restart may stop its listener after shutDown notified its listeners, or after the next instance became current. Its drain then ends the connections with the shutdown notice and registers with no server. - A drain whose last connection closed while it registered deregisters itself, so a finished drain no longer stays a shutdown listener. - Tests: the in-core restart case reads the notice while the server restarts, since macOS drops unread data of a loopback connection closed for longer than net.inet.tcp.fin_timeout. New case: a drain of a handler whose server is shutting down while another instance is current. The queued-notice case checks that the drain ends within a second once its connection is closed.
70ec861 to
5aa486c
Compare
|
Thanks for the second round. Points 1–5 are taken in 5aa486c, and point 6 is left as you suggested. The branch is rebased onto master 3295ece (#1123) without conflicts. 1. The restart case on macOS. Taken as you wrote it: the notice is read on another thread while 2. A drain started after 3. A finished drain left registered. Taken as suggested: after registering, a drain that is already finished deregisters itself. 4. The close half of 5. The restart case does not force the order. New case 6. The catch arm in Checks.
The PR description is updated to match. |
…t a plain bind accepts The two cases whose handler has allow-tcp-reuse-address: false failed on the ubuntu/JDK 11 leg of both CI runs of OpenIdentityPlatform#1125: the port check of initializeConnectionHandler could not bind 127.0.0.1:6552x. The port came from TestCaseUtils.findFreePort(), which checks with SO_REUSEADDR only, and every test class counts ports down from 65535 in a JVM of its own, so a port can still carry a connection an earlier class left in TIME_WAIT. Such a port accepts a bind with SO_REUSEADDR and refuses one without it. These cases now take a port that a plain bind on 127.0.0.1 accepts, and skip the others.
|
CI: the red ubuntu/JDK 11 leg. Fixed in a96d796. It is a test fixture problem, not a problem in the handler. The leg failed in both CI runs of this PR, on 70ec861 and on 5aa486c. Each time the same two cases of The cause is in the ports. The two cases now take their port from a small helper, Checks. The class was run directly with TestNG while another process held ports 65505–65533 in the state CI met: each port had an open connection and no listener, so a bind with |
maximthomas
left a comment
There was a problem hiding this comment.
praise: The drain now decides on the server it belongs to, and the tightened tests measurably pin the drain.
TransportDrain.start()testsserver, the instance captured ininitializeConnectionHandler(LDAPConnectionHandler2.java:681,:206,:215). A drain stopped during an in-core restart therefore never registers with the next instance and never waits for a listener loop that has already run.serverShutdownDeliversANoticeQueuedBehindUnreadDatanow joins forCLOSED_DRAIN_ENDS_MS. EmptyingOpenConnections.onCloseEventturns it red ("the drain is still waiting for a closed connection", 1 of 19 cases red), so the close half of the drain is pinned.serverRestartEndsTheConnectionsOfAListeningHandlerreads the notice while the server restarts. The class is green 19/19 locally on macOS and on every Linux CI cell.
Fixes #1119
Summary
LDAPConnectionHandler2, the class every shipped LDAP, LDAPS and administration listener runs on, bound its listeners to the transport the SDK shares across the JVM. Souse-tcp-keep-alive,use-tcp-no-delay,buffer-sizeandnum-request-handlershad no effect, andallow-tcp-reuse-addressonly reached the "address in use" pre-check. Each handler now gets a transport of its own, built from its configuration, asHTTPConnectionHandleralready does.Changes
SDK:
GrizzlyLDAPListenerGRIZZLY_TRANSPORT, modelled onGrizzlyLDAPConnectionFactory.GRIZZLY_TRANSPORT. The listener binds to that transport and serves its connections with it. It does not shut the transport down when it is closed: the owner stays responsible for it. Without the option, nothing changes.LDAPConnectionHandler2TCPNIOTransport:num-request-handlerssets the number of selector threads. When it is unset, the system propertyorg.forgerock.opendj.transport.selectorsis used, so JVMs tuned with it keep their setting. Without the property, the count ismax(2, cpus/2), as for the legacy and HTTP handlers.allow-tcp-reuse-addresssetsSO_REUSEADDRon the listen socket.buffer-sizesets the write buffer only, matching its description ("LDAP response message write buffer") andADMIN_WRITE_BUFFER_SIZE. It is deliberately not the read buffer: every shipped handler hasds-cfg-buffer-size: 4096 bytesexplicitly, and 4 KiB reads would be a regression. Today reads are bounded bySO_RCVBUF.use-tcp-keep-aliveanduse-tcp-no-delayare passed asLDAPListener.SO_KEEPALIVEandTCP_NO_DELAY.LDAPServerFilterapplies the listener options to every accepted socket after the transport has applied its own, so the transport cannot carry them.PooledMemoryManager. It pre-allocates 3% of the heap as direct buffers, so one per transport would multiply that by the number of handlers. Memory use stays what the shared transport had.SERVER_SHUTDOWNwith a notice of disconnection, then the transport is shut down. The server does not disconnect clients itself on shutdown: that used to happen when the last listener released the shared transport.shutdownNow()would drop it. A client that does not read within those 2 seconds still gets EOF.Configuration:
LDAPConnectionHandlerConfiguration.xmluse-tcp-keep-aliveanduse-tcp-no-delayare now defined locally instead of referenced fromPackage.xml, and requirecomponent-restart;buffer-sizerequires it too. The transport fixes these values when the listener starts, likeaccept-backlog,max-request-size,num-request-handlersandallow-tcp-reuse-address, which already required it.Package.xmlis not touched: the HTTP connection handler and the LDAP pass-through policy share those definitions and apply them.opendj-server-legacynow declares its dependency onopendj-grizzly; it used to come in only throughopendj-server.Behaviour changes worth noting
max(5, cpus/2 - 1)threads, each handler has its own. With the defaults:max(2, cpus/2)for LDAP and for LDAPS, and 4 for the administration connector.use-tcp-keep-alive,use-tcp-no-delayorbuffer-sizeon an LDAP connection handler now reports that a restart of the handler is required. The XML is shared by both classes, and ADM cannot set the requirement per class. So the same is reported for a handler that names the legacyorg.opends.server.protocols.ldap.LDAPConnectionHandlerexplicitly, although that class applies these values to the next accepted connection. Since dsconfig create-connection-handler --type ldap creates the legacy LDAPConnectionHandler, not the LDAPConnectionHandler2 the server ships with #1116 new handlers default toLDAPConnectionHandler2.buffer-sizeof 4096, a response larger than about 6 KiB is now written in severalwrite()calls. Grizzly caps a write at 3/2 of the write buffer. Before, the cap was the size of the socket send buffer. No benchmark has been run.Tests
New
LDAPConnectionHandler2TransportTestCase(19 cases) runs a real handler built from a configuration entry and inspects the server-side sockets:SO_REUSEADDRon the listen socket, true and false. The cases whose handler has noSO_REUSEADDRtake a port that a bind without it accepts on 127.0.0.1:TestCaseUtils.findFreePort()checks withSO_REUSEADDRonly, and every test class counts ports down from 65535, so a port can still carry a connection an earlier class left inTIME_WAIT;num-request-handlers, from the system property, from the processor count, andnum-request-handlerstaking precedence over the property. Each case also checks that stopping the handler, with no connection open, stops its threads;UNAVAILABLE) and the shutdown reason. The drain is registered as a shutdown listener;GrizzlyLDAPListenerTestCase#testLDAPListenerWithProvidedTransport: the listener serves its connections on the given transport and leaves it running when closed.Reactor run (
-pl opendj-server-legacy -am, precommit):RejectedSSLConfigurationChangeTestCase24/24.LDAPConnectionHandler2TransportTestCaserun directly with TestNG: 19/19, three runs in a row. In round 3, with ports 65505–65533 held the way CI left them (a bind withSO_REUSEADDRpasses, one without it fails), the round-2 test class fails the same two cases as the ubuntu/JDK 11 leg, on the same ports, and the round-3 one passes 19/19; it also passes 19/19 twice on free ports.GrizzlyLDAPListenerTestCase11/11 in the first round; the SDK has not changed since.Mutation check, first round: 14 of 14 mutants fail the new test class. Each removes one piece:
GRIZZLY_TRANSPORT;GRIZZLY_TRANSPORT;Mutation check, review round 1: 11 of 13 mutants fail the class:
num-request-handlers;The two survivors were the late
isShuttingDown()read and deregistering a drain that was never registered.Mutation check, review round 2: 13 of 16 mutants fail the class. The late
isShuttingDown()read is now pinned: checking the current instance instead of the handler's fails the new case, and so does dropping the check after registration. The three survivors each differ only inside a window a test cannot hit on demand:getNewInstanceandbootstrapServer;Related: #1116 / #1124 make
dsconfig create-connection-handler --type ldapcreate handlers onLDAPConnectionHandler2, so this fix also covers the handlers administrators create.