[SPARK-58021][CONNECT] Add forceful local pool purge - #58248
Open
ericm-db wants to merge 4 commits into
Open
Conversation
This was referenced Aug 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
This is layer 7 of the nine-PR local Connect pool stack:
#57684 -> #57685 -> #57907 -> #57686 -> #58247 -> #57687 -> #58248 -> #57102 -> #57688
Until lower layers merge, GitHub shows their cumulative diff. The review unit introduced here is
commit
196221a8d5a.This layer adds the forceful escape hatch for returning the pool to a clean slate:
python -m pyspark.sql.connect.local_server_pool --purge.SparkSession integration and JIT warmup remain in later PRs.
Why are the changes needed?
Normal retirement deliberately preserves retryable state. Operators also need a bounded,
destructive recovery path when the pool itself is corrupt or wedged. Keeping purge separate makes
its stronger signalling and deletion semantics explicit and independently reviewable.
Does this PR introduce any user-facing change?
Yes. It adds
python -m pyspark.sql.connect.local_server_pool --purge, which force-stops everylocal pool process it can verify and empties the pool directory. SparkSession still does not select
the pool until #57102.
How was this patch tested?
Added one focused corruption-and-process-lifecycle test at this layer, bringing the suite to 51
tests. It covers ready, pending, half-started, duplicate-claimed, retiring, and malformed members.
At the stack tip, the equivalent direct
unittestinvocation passed all 56 pool tests, includingthe two real-server E2E tests.
The rebuilt commit passed Python AST parsing,
git diff --check, and changed-file ASCII and100-column checks.
Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude Code (Fable 5) and OpenAI Codex (GPT-5)