Skip to content

fix: drop held feedback across a reconnect instead of resending it - #5

Merged
tactino merged 1 commit into
mainfrom
fix/drop-feedback-across-reconnect
Sep 11, 2026
Merged

tactino merged 1 commit into
mainfrom
fix/drop-feedback-across-reconnect

Conversation

@tactino

@tactino tactino commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

The bug

feedback() retried its send on a new connection after every close except an explicit resync. That reads as resilience. It is the opposite.

A transition spans two messages. The infer the server answered is where it recorded the previous observation, the policy step state and the done flags — and all of that lives inside the connection handler. A new connection starts with none of it.

So a resent feedback does not save the step. The server completes the transition from an empty observation and stores it. No exception, no warning, one corrupt transition in the training buffer per reconnect.

How it was found, and one correction

A clean-machine run of the documented quickstart dropped its connection and reconnected. The first diagnosis blamed WebSocket keepalive pings colliding with a CPU-bound learn step; that was then measured and is false — plugrl-server runs learn off the event loop, and learns of 190 s produce no ping timeout. The actual cause was the machine suspending for nearly two hours mid-run.

That correction does not touch this pull request. The state-loss path was read out of the server's handler and reproduced in a unit test, not inferred from the incident. Only the cause was inferred, and the cause is the part that was wrong. Every reconnect costs the same state, whatever produced it: a suspend, a flaky link, an operator restarting the server.

The measurement is written up in PlugRL/plugrl-server#4 as experiments/e8-keepalive-hypothesis/.

E6, the learning-curve experiment, was checked against this: one connection per seed, zero reconnects, so that data is unaffected.

The change

  • every close in the feedback path now drops the transition and returns;
  • the next infer reconnects and resynchronises, which is what the exchange beginning with an infer is for;
  • the infer path deliberately still retries, and a test asserts that, so a later cleanup does not make both paths symmetric.

The resync branch already did the right thing. It was simply applied to the rarest close reason instead of all of them.

Tests

tests/test_reconnect_drops_feedback.py, 4 cases: keepalive timeout, plain close, resync, and the infer mirror image. Verified to fail against the previous behaviour rather than passing vacuously. Full suite 135 passed, 10 skipped. ruff check --exclude third_party . clean.

Specified as SPEC §7.6 in PlugRL/plugrl-protocol#3.

🤖 Generated with Claude Code

The feedback path retried on a new connection after every close except an
explicit resync. That looks like resilience and is the opposite.

A transition is built across two messages. The infer the server answered is
where it recorded the previous observation, the policy step state and the
done flags, and all of that lives in the connection handler. A new connection
starts with none of it. So a feedback resent after a reconnect does not save
the step: the server completes the transition from an empty observation and
stores it, and nothing anywhere reports an error. The buffer just quietly
gets a hole in it, and gets more of them the more often the connection drops.

The resync branch already knew this and returned. It was the only branch that
did, and resync is not the common way a connection dies - a keepalive ping
timeout is, which is what turned this up. So every close now drops the
transition and lets the next infer resynchronise, which is the whole point of
the exchange starting with an infer.

The infer path deliberately still retries. A fresh infer on a fresh
connection is exactly how the two sides get back in step, and the tests
assert that too, so a later cleanup does not "fix" both paths the same way.

Specified as SPEC section 7.6 in PlugRL/plugrl-protocol#3; the server half,
which stops provoking the reconnect in the first place, is
PlugRL/plugrl-server#4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@tactino
tactino merged commit 089faca into main Sep 11, 2026
2 checks passed
@tactino
tactino deleted the fix/drop-feedback-across-reconnect branch September 11, 2026 19:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant