Skip to content

fix(utilities): stop truncating every traceback for a loud BaseError - #983

Merged
JarryShaw merged 2 commits into
mainfrom
fix-719-tracebacklimit-leak
Oct 2, 2026
Merged

JarryShaw merged 2 commits into
mainfrom
fix-719-tracebacklimit-leak

Conversation

@JarryShaw

@JarryShaw JarryShaw commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What is the purpose of your pull request?

  • fix — corrects a defect

Description of your pull request and other information

Issue #719: a loud BaseError set sys.tracebacklimit = 0 outside development mode and never restored it. That attribute is process-global, so one pcapkit error silently truncated every subsequent traceback in the process, including exceptions pcapkit had nothing to do with. Measured: an unrelated IndexError raised after a loud IntError came back with traceback.extract_tb() returning 0 frames.

The fix removes the assignment outright. The same terse output now comes from exception hooks installed lazily on sys.excepthook and threading.excepthook, the first time a loud error outside development mode actually needs them. A hook prints limit=0 against the exception's real traceback, not tb=None — the latter only discards the top exception's own frames and leaves a chained cause's traceback printed in full; fixed after cross-review caught it, with tests covering an explicit raise ... from err and an implicit context chain. Everything else is delegated, unchanged, to whatever hook was previously installed. Install is idempotent by identity against this module instance's own hook function, not a marker shared across every reloaded copy of it — the first version of this fix used the marker and CI caught the difference under pytest -n auto. A re-entrant call falls back to the interpreter's default instead of recursing; the thread-side guard is per-thread (threading.local), since two worker threads can legitimately raise concurrently without being each other's re-entrancy.

I now set threading.excepthook, reversing an earlier decision to decline it. The interpreter's default thread hook itself consults sys.tracebacklimit, so a loud BaseError on a worker thread used to be terse only as a side effect of the global this PR removes — leaving it unset regressed that path from 1 line to 12 (measured), which my first write-up here wrongly called acceptable. That gap is now closed, including the thread-identifying header line the default hook always prints first.

tests/utilities/test_decorators.py gained a tearDown this leak itself needed: constructing a loud StructError there with no quiet=True was already leaving sys.tracebacklimit = 0 set for every later test in the process on current main — the same #981 shape, just not previously visible as a failure.

Remaining limitation: neither hook fires for a traceback a caller formats itself with traceback.* instead of letting it reach the top level uncaught. Strictly better than today regardless, since sys.tracebacklimit truncated that path too, for every caller in the process.

@JarryShaw JarryShaw added fix Pull requests that fix a defect (fix: subject prefix) review: pending No verdict for the current head - never reviewed, or the head moved since the last one test Pull requests that add or correct tests (test: subject prefix) labels Oct 1, 2026
@JarryShaw

Copy link
Copy Markdown
Owner Author

A premise of my own brief was wrong, and it is worth correcting before this is reviewed.

When I recommended against the excepthook I said an exception hook installed merely by importing the
library is a side effect on the host, and I briefed the worker to install lazily on first use to avoid
it. pcapkit already installs one eagerly at import. pcapkit/__init__.py:86-94, outside
development mode:

ROOT = os.path.dirname(os.path.realpath(__file__))
tbtrim.set_trim_rule(lambda filename: ROOT in os.path.realpath(filename),
                     exception=BaseError, strict=False)

tbtrim replaces sys.excepthook and threading.excepthook at import, and deliberately exempts
BaseError from its trim rule. So the objection I raised against option (1) was already moot — a host
application's own hook is clobbered by import pcapkit regardless of what we do in exceptions.py.
Lazy installation is still the better behaviour and is what landed, but it is a smaller improvement
than I implied.

A second finding, independent of this change. tests/utilities/test_decorators.py already leaks
sys.tracebacklimit on current main — it builds a loud StructError with no quiet=True and has no
teardown. Measured:

before False
after  True 0        <- 22 tests pass, and the global is left set

That is #981's shape again, in a second file, and it broke the new test under a full-directory run
until the worker added an addCleanup. So the invisible-global-leak class has now bitten twice,
which strengthens the case for whichever fix #981 takes.

Also now stale, and not this change's to fix: pcapkit/corekit/enum.py:361-370 still describes the
old "would set sys.tracebacklimit" behaviour. That file belongs to the citation change in flight; I
will pick it up after that lands rather than create a conflict.

Labelled fix, test, review: pending. Cross-review going out on a different model.

@JarryShaw
JarryShaw force-pushed the fix-719-tracebacklimit-leak branch from cd1a271 to 3f1d505 Compare October 2, 2026 00:02
@JarryShaw

Copy link
Copy Markdown
Owner Author

NEEDS CHANGES at cd1a27168 — opus cross-review. Two blocking findings, both of which I
reproduced. One of them refutes a sentence I relayed on this PR, so that correction comes first.

1. The hook does not reproduce tracebacklimit = 0 — a chained exception still prints the cause's
frames, including ours.
tb=None suppresses only the top exception's traceback; __cause__ keeps
its own, where the global flattened the whole chain. Measured on a wrapped ZeroDivisionError:

tracebacklimit = 0            ZeroDivisionError: division by zero
                              The above exception was the direct cause of ...
                              E: wrapped
print_exception(t, v, None)   Traceback (most recent call last):
                                File ".../chain.py", line 8, in inner
                              ZeroDivisionError: division by zero
                              ... + the rest
print_exception(t, v, tb, 0)  byte-identical to tracebacklimit = 0

This is the common shape, not a corner: 42 raise-inside-except sites across 22 files under
pcapkit/ by an AST scan — my own count, higher than the review's 36. The fix is the fourth argument.
The test suite has no chained case, which is why it got through.

2. A loud error on a worker thread regresses from one line to twelve, and my earlier "strictly better
than today" was wrong for that path.
The default thread hook prints with limit=None, which consulted
sys.tracebacklimit — so the global was what made the thread path terse. Measured:

with tracebacklimit = 0   2 lines
without it                12 lines

I repeated the PR body's framing when I labelled this PR. It is accurate for traceback.* calls a caller
makes itself; it is false for the thread path, which gets strictly more verbose. Either set
threading.excepthook with the same capture-and-chain treatment, or say so plainly in the body and
docstring rather than calling it an improvement.

Also found, non-blocking. One new test is vacuous — ChainingTests::…delegated_to_the_previously _installed_hook passes on unfixed code, because with no hook installed the script's own previous_hook
is sys.excepthook. And four places outside the five changed files still assert the removed mechanism
as current behaviour, including the module's own user-facing page at
docs/source/pcapkit/utilities/exceptions.rst:16,45 and pcapkit/corekit/enum.py:367.

One fact worth recording. tbtrim is configured with exception=BaseError, which exempts our
errors from its trimming — so the old terseness for a loud BaseError came entirely from
sys.tracebacklimit, not from tbtrim. That is why removing it has more reach than it appears to.

Verified sound: the simple uncaught case is byte-identical to before; devmode unchanged; the quiet path
fully inert; a host hook installed after import now receives the real traceback where it was crippled
before; a raising delegate swallows nothing; and the re-entrancy guard replaces a RecursionError.

@JarryShaw

Copy link
Copy Markdown
Owner Author

I think we can refer to tbtrim to see how it's being implemented and if we'd still have to build the wheel.

@JarryShaw

Copy link
Copy Markdown
Owner Author

NEEDS CHANGES at 3f1d5057c — opus cross-review, round 2. Five items. The guard fix itself is
sound: same-instance idempotency holds, two loud errors leave the hook stable, and
_previous_excepthook is _excepthook stays false.

1. Both blocking findings survive this head, measured here. Chained exceptions still print the
cause's frames — 12 non-blank stderr lines against 3 on the pre-fix tree, with no chained-case
test anywhere in tests/utilities/. The thread path still regresses — 13 lines against 1
pre-fix. Its AST scan matched mine exactly at 42 raise-in-except sites across 22 files, of which
32 raise a pcapkit.utilities.exceptions class.

2. The description now contradicts its own code. The body says "Installing twice, or from a
reloaded copy of the module, is a no-op rather than wrapping itself."
The reloaded-copy half is
precisely the bug the amend removed
— the old marker-based guard did treat a reloaded copy as a
no-op, which is why CI went red. It must now install over it.

3. A test proves something the real world cannot produce. The docstring names "a genuine
importlib.reload" as the scenario where instance B installs over A's hook. That path cannot occur:
pcapkit/__init__.py:93 calls tbtrim.set_trim_rule, and tbtrim:83 assigns sys.excepthook
unconditionally — I confirmed both. So a fresh package import evicts A's hook before B installs,
and B's _previous_excepthook is a new tbtrim_excepthook, not A's. CrossInstanceReinstallTests
passes only because the harness loads modules by spec and never executes pcapkit/__init__.py. Not a
defect in the fix, but the docstring claims a guarantee for a path that does not exist.

4. The vacuous test is still vacuous. It ran the test's verbatim body against the pre-fix tree:
all three assertions pass, including the new assertNotIn. Nothing asserts the hook is installed.

5. The stale prose was not taken — the diffstat is still 5 files, none under docs/, and
test_stacklevel.py is untouched.

Test hygiene passes, and the check was the right one: 64 tests across the touched modules run twice
in one interpreter, with the state diff after round 2 empty. Round-1 drift is first-import effects
only and does not accumulate; tracebacklimit stayed unset and no logger is left disabled.

It also corrected round 1's reload nit: in a genuine reload tbtrim's trimming is not lost, because
the fresh instance captures tbtrim's fresh hook. That nit bites only in the harness path.

UNVERIFIED by it: the full directory double-run (it substituted 64 tests after the first attempt
overran), the pytest runner, the lint checklist, and CI on this head.

@JarryShaw

Copy link
Copy Markdown
Owner Author

Looked at it, and tbtrim cannot produce the one-line output — but what it does instead may be what you
actually want.
Measured against main, re-arming tbtrim without the exception=BaseError exemption:

                         plain loud error   chained
exception=BaseError            1 line        3 lines   <- today, terse via sys.tracebacklimit
no exemption (tbtrim trims)    7 lines      14 lines

Why dropping the exemption does not give terse output. tbtrim.py:66-81 counts frames until the
first file matching the predicate, then prints with that limit. For an error raised from user code
the outermost frames are the caller's, so the limit is non-zero and tbtrim prints the user's frames and
hides ours
. The limit == 0 branch — the one-line form — only fires when the whole stack is already
inside pcapkit. So today's terseness came from sys.tracebacklimit, never from tbtrim, which is why
removing the global broke it.

Two things worth knowing before deciding. tbtrim.py:78-79 uses the very same
print_exception(etype, value, None) idiom this pull request was just faulted for, so reusing tbtrim
would inherit the identical chained-exception shortfall
— the cause's frames survive either way. And
tbtrim does wire threading.excepthook at :84-87, so delegating to it is the cheap way to fix the
thread regression without hand-rolling a second hook.

That opens a third option I had not put to you, and I think it is worth a moment's thought rather than
a quick no.
Instead of one line, let tbtrim trim: a reader sees their own frames and none of ours.
That is arguably the more useful output — it tells them where their call went wrong rather than
withholding the location entirely — it needs no new hook, no new global, and no capture-and-chain logic,
and it is one line removed from pcapkit/__init__.py:93.

The cost is that it is not the behaviour you asked to preserve. One line becomes about seven, and
"handled output" would come to mean trimmed rather than terse.

No change made. If you want the terse form kept, the current approach stands and the worker fixes
the chained and thread paths. Say which and I will take it.

@JarryShaw JarryShaw added the needs: decision Waiting on the maintainer to decide — not blocked by other work label Oct 2, 2026
…719)

A loud BaseError set sys.tracebacklimit = 0 outside development mode and never
restored it. The attribute is process-global, so one pcapkit error silently
truncated every subsequent traceback for the rest of the process, including
exceptions pcapkit had nothing to do with. Measured: an unrelated IndexError
raised after a loud IntError came back with traceback.extract_tb() returning
0 frames.

Removed the assignment outright. The same terse output for a loud error now
comes from exception hooks installed lazily on sys.excepthook and
threading.excepthook, the first time a loud error outside development mode
actually needs them -- never at import time. A hook prints limit=0 against
the exception's real traceback, not tb=None, which only discards the top
exception's own frames and leaves a chained cause's traceback printed in
full -- cross-review caught this; tests now cover an explicit `raise ... from
err` and an implicit context chain. Everything else is delegated, unchanged,
to whatever hook was previously installed. Install is idempotent by identity
against this module instance's own hook function, not a marker shared across
every reloaded copy of it -- CI caught the marker version under
`pytest -n auto`. A re-entrant call falls back to the interpreter's default;
the thread-side guard is per-thread (threading.local), since two worker
threads can legitimately raise concurrently without being each other's
re-entrancy.

threading.excepthook is now set, reversing an earlier decision to decline it.
The interpreter's default thread hook consults sys.tracebacklimit itself, so
a loud BaseError on a worker thread was terse only as a side effect of the
global this removes; leaving it unset regressed that path from 1 line to 12
(measured). That gap is closed, including the thread-identifying header line
the default hook always prints first.

tests/utilities/test_decorators.py gained a tearDown this leak itself needed:
constructing a loud StructError there with no quiet=True already left
sys.tracebacklimit = 0 set for every later test in the process on current
main -- the same #981 shape, just not previously visible as a failure.

Development mode is unchanged. Neither hook fires for a traceback a caller
formats itself with traceback.* rather than letting it reach the top level
uncaught -- strictly better than today regardless, since sys.tracebacklimit
truncated that path too, for every caller in the process.
@JarryShaw
JarryShaw force-pushed the fix-719-tracebacklimit-leak branch from 3f1d505 to afa7435 Compare October 2, 2026 00:33
@JarryShaw

Copy link
Copy Markdown
Owner Author

All five round-2 items addressed at afa7435ac, and the worker reversed its own position on the one I
left to its judgement.
This does not clear the needs: decision above — if you take the trimmed
option, most of this is discarded. Flagging that plainly so nothing here reads as a decision already made.

Chained exceptions fixed and now tested. print_exception(etype, value, tb, 0) — the real traceback
with limit=0 — confirmed byte-identical to the old mechanism for explicit and implicit chains. Two new
tests compare against the old shape rather than asserting "one line", and both fail with the exact
mismatch when reverted to tb=None.

It reversed itself on threading.excepthook after measuring, and the reasoning is better than mine.
It confirmed 2 lines versus 12 itself, then set the hook with the same capture-and-chain design — with two
adaptations it worked out rather than copied: ExceptHookArgs arrives as one argument, not three, and the
re-entrancy guard is a threading.local() rather than the shared flag used for sys.excepthook.
It proved that distinction empirically: swapping in shared state made the concurrency test hang with
1 of 2 threads completing, because one thread's hook mistook another's concurrent call for its own
re-entrancy. It also found its first attempt silently dropped the "Exception in thread {name}:" header
the default hook always prints, and restored it.

It then edited the PR body itself to narrow the "strictly better than today" claim to caller-formatted
traceback.* output, rather than leaving a self-contradicting description live. That was the right
instinct — the body is an artifact, not a note.

The vacuous test now bites. It asserts the hook is installed before checking delegation, and against
the pre-#719 tree it fails with INSTALLED: False.

Stale prose taken where I scoped it — docs/source/pcapkit/utilities/exceptions.rst and
tests/utilities/test_stacklevel.py are now past tense and reference #719; pcapkit/corekit/enum.py and
registry-protocol.rst left for the change that owns them.

Verified: tests/utilities/ 121 passed under pytest and 121 OK under plain unittest; pylint
10.00/10, mypy clean, isort clean; every new assertion proven to fail without its own fix by targeted
revert. 7 files, +1147/−49.

Still gated on your terse-versus-trimmed call. A fresh cross-review goes out once that is settled —
reviewing this head before the direction is fixed risks spending a round on code that may not survive.

@JarryShaw

Copy link
Copy Markdown
Owner Author

The red mark on afa7435ac is a runner eviction, not a defect — but my check for an in-log test
failure was inconclusive, so I am not calling the run clean.

step 6  "Run unit tests"  -> cancelled
##[error]The runner has received a shutdown signal...
##[error]The operation was canceled.
progress reached [ 81%]
run totals: 48 success, 2 failure, 1 skipped

The two failures are Python 3.12 and the aggregate Required checks passed, so one cause, two
marks
. The step did not fail — it was cancelled mid-run by the runner going away at 81%, which is
the same shape as the apt-mirror stall that produced #974.

What I could not establish: whether any test had already failed before the cancellation. The
progress output is dots and u markers, and the pattern I grepped for an F was loose enough to match
other text, so I will not claim the run was green up to that point. It needs a clean re-run to know —
and I will not trigger one, since re-running workflows is not mine to do.

This does not change the PR's state: it is still held at needs: decision on the terse-versus-trimmed
question, and the local evidence stands on its own — tests/utilities/ was 121 passed under pytest
and 121 OK under plain unittest at this head, with pylint 10.00/10, mypy and isort clean.

@JarryShaw

Copy link
Copy Markdown
Owner Author

What exactly needs my decision and what's your recommendation?

@JarryShaw

Copy link
Copy Markdown
Owner Author

One decision here, and my recommendation is: change nothing — keep the terse output.

The question. exceptions.py:201 used to set sys.tracebacklimit = 0 for every loud BaseError outside
DEVMODE and never restore it, so the library mutated a process-global that affects all tracebacks, not just
ours. This pull request replaces that with a lazily-installed sys.excepthook plus a threading.excepthook,
keeping the one-line output. The alternative I surfaced was to drop the exception=BaseError exemption at
pcapkit/__init__.py:93 and let tbtrim trim instead — one line removed, no new hook, no capture-and-chain
logic.

Why I recommend against it. Measured against main:

                         plain loud error   chained
exception=BaseError            1 line        3 lines   <- today
no exemption (tbtrim trims)    7 lines      14 lines

tbtrim.py:66-81 counts frames to the first file matching its predicate and prints with that limit, so for
an error raised from user code the limit is non-zero and it prints the caller's frames while hiding ours.
The limit == 0 one-line form only fires when the whole stack is already inside pcapkit. So trimmed is not
a cheaper way to get terse output — it is different output.
You ruled at 22:53Z that you prefer the output
handled; trimmed would overturn that, not implement it.

Two further things against it: tbtrim.py:78-79 uses the same print_exception(etype, value, None) idiom this
pull request was faulted for, so reusing tbtrim inherits the identical chained-exception shortfall; and
taking it would discard a change that is now complete and verified — tests/utilities/ 121 passed under
pytest and 121 OK under plain unittest, pylint 10.00/10, mypy and isort clean, with every new assertion
proven to fail without its own fix by targeted revert.

I raised trimmed only because I had not put it to you, and I am not advocating it. If you say "keep terse"
I will drop needs: decision, dispatch the cross-review that I have been holding back, and this moves like
the others.

One thing you should know before deciding, because it is not clean. The two red marks on afa7435ac are a
runner eviction, not a defect — step 6 was cancelled at 81% by The runner has received a shutdown signal, surfacing as Python 3.12 plus the aggregate, so one cause and two marks. But I could not
establish whether any test had already failed before the cancellation: the progress output is dots and u
markers, and the pattern I grepped was loose enough to match other text. It needs a clean re-run to know, and
re-running workflows is not mine to do. The local evidence above stands on its own.

The other three decisions are on #982, with recommendations.

@JarryShaw

Copy link
Copy Markdown
Owner Author

keep terse

@JarryShaw JarryShaw added review: good-to-go Cross-review at the current head says ready; CI state is separate and removed needs: decision Waiting on the maintainer to decide — not blocked by other work review: pending No verdict for the current head - never reviewed, or the head moved since the last one labels Oct 2, 2026
@JarryShaw

Copy link
Copy Markdown
Owner Author

GOOD TO GO at afa7435ac — sonnet cross-review, held until your terse ruling and now run in full.
Labelled review: good-to-go. It also settled the thing I twice told you I could not establish, and
corrected a claim I relayed.

The red marks: nothing had failed before the eviction. I re-derived this myself from the job log —
zero FAILED/ERROR lines, zero "N failed" summary text, and the progress characters are only . and
u, with no F or E anywhere before ##[error]The runner has received a shutdown signal at 00:54:23Z,
following progress at 81%. Every Compat leg (3.10-3.14) and every Engines leg (3.10-3.14 × 6 backends) is
green. One honest limit: I still cannot say what u denotes. pytest-xdist is not installed in the repo
venv and I will not install into it, and the repo's conftest does not customise the status character. So the
conclusion rests on the absence of failure markers and summary text, not on decoding u.

A correction to what I relayed earlier. I told you the worker "proved it empirically: swapping in shared
state made the concurrency test hang with 1 of 2 threads completing". The misattribution bug is real and
independently confirmed
— with a shared bool and deterministic overlap, thread 2's own unrelated
BaseError is misdiagnosed as thread 1's re-entrant call and falls through to a full traceback instead of
the terse one. But the hang did not reproduce. The reviewer's own first apparent hang came from a
synchronisation bug in its harness, not from the shared-flag design. So threading.local() is right for a
confirmed correctness reason; the hang was either a stuck test misread or unreproducible. Either way the
shipped design stands.

The install guard is load-bearing, proven by revert. Reverting to the old marker form
(getattr(current, 'installed_by_pcapkit', False)) makes
CrossInstanceReinstallTests::test_a_fresh_instance_installs_over_a_stale_instance_hook fail exactly as the
description claims. Self-wrapping is unreachable by construction, a double install is a no-op, and a third
party's hook installed between two of ours is captured and chained.

The tbtrim composition resolves correctly, against the real thing. tbtrim assigns both hooks
unconditionally at import pcapkit; the first loud BaseError then captures tbtrim's hook as the delegate
and installs over it. LoudErrorPrintsOneLineTests proves terseness against a genuine import pcapkit in a
subprocess — not the by-spec harness load that skips __init__.py, which a previous round correctly
warned about and which is used only by the narrower TracebacklimitNeverSetTests.

No process-global residue: two rounds in one interpreter, sys.tracebacklimit unset both times, both
hooks the same object across rounds, 39 passed each round; tests/utilities/ 121 passed / 104 subtests.

One real gap it names, not a blocker: no test asserts end-to-end that a non-BaseError exception is
trimmed by the real tbtrim hook after a loud error — ChainingTests uses a user stub. Signatures match by
inspection. Worth a follow-up test.

I am not calling this ready to merge, because the two red marks are still unresolved on the revision
itself and clearing them needs a re-run that is not mine to trigger. The verdict and the local evidence are
in; the CI state is yours to judge.

UNVERIFIED: the u marker's meaning; a sample of the 42 real raise-inside-except sites; revert-tests for the
thread-hook and double-install tests specifically.

@JarryShaw

Copy link
Copy Markdown
Owner Author

Unit tests failed?

@JarryShaw

Copy link
Copy Markdown
Owner Author

READY TO MERGE at 3c0474391. review: good-to-go, CI 59 ok / 0 fail / 0 incomplete, merge state
CLEAN.

Your merge of main into the branch resolved the one thing I could not. The two red marks on afa7435ac
were a runner eviction that I could only characterise, never clear — a re-run was not mine to trigger. Merging
main produced a fresh run, and it is green. The eviction question is closed by a measurement rather than by
my argument about it.

The verdict carries to the new head, and here is why rather than just the assertion. 3c0474391 is a
merge commit; afa7435ac is an ancestor of it (git merge-base --is-ancestor confirms), and the only content
delta is tests/vendor/test_vendor_snapshot_restore_unit.py — #986's fix, which was separately reviewed,
labelled good-to-go, and has already merged as 6f41995b8. So nothing unreviewed entered this branch.

Unpublished and otherwise untouched. Yours.

@JarryShaw
JarryShaw merged commit 390528a into main Oct 2, 2026
63 checks passed
@JarryShaw
JarryShaw deleted the fix-719-tracebacklimit-leak branch October 2, 2026 04:05
@JarryShaw JarryShaw removed the review: good-to-go Cross-review at the current head says ready; CI state is separate label Oct 2, 2026
JarryShaw added a commit that referenced this pull request Oct 3, 2026
…issue (#719)

- Nine citations in tests/ called a pull request "GitHub issue"
  (#921, #764, #906, #815, #936, #721, #501, #983) or lumped PR #428 in
  with issue #425; each now names the right kind.
- test_dispatch_default_resolution_unit: issue #425 reported the registry
  leak and PR #428 fixed it, so the sentence says "reported" and "fixed"
  instead of crediting the issue with the fix.
- Prose only: docstrings and comments, no assertion or logic touched.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fix Pull requests that fix a defect (fix: subject prefix) test Pull requests that add or correct tests (test: subject prefix)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant