perf: accelerate exact count-table parsing and preserve strict input boundaries - #149
Merged
Merged
Conversation
dnncha
marked this pull request as ready for review
October 4, 2026 08:45
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ordinary raw-count cells previously passed through regex and Decimal parsing. Parse unsigned ASCII integers of at most 64 characters directly as exact integers, preserving counts above 2**53. Longer inputs, signed values, decimals and scientific notation retain the existing Decimal path and limits; Unicode digits remain invalid and unselected sample columns remain validated.
The production diff against current main is the five-line count-parser fast path. Main already contains the equivalent FASTQ printable-ASCII optimization, and this branch retains main's implementation. Added 537 boundary cases cover byte-range characters, Unicode, lossless counts, conversion limits and unselected columns.
The reproducible benchmark and raw evidence in
docs/audits/2026-09-30-input-throughput.{md,json}record the original 30 September baseline/candidate pair. Five alternating Linux/CPython 3.12 runs measured count-table parsing at 0.468 → 0.176 seconds (2.66x throughput) for 10,000 guides × 32 samples, with identical complete parsed-output hashes. These are synthetic Python parsing measurements, including output hashing; they do not establish whole-assay speedups. The historical FASTQ measurement describes the optimization already integrated into main. These measurements were not rerun during reconciliation.Current integration validation incorporates main through
2fa3532549d5945991aba2d7645af66108206497, including the strict native gzip fix:make all shared test cli-test python-test: native build, both C test executables, CLI fixtures and native/Python FASTQ parity pass; 2,069 Python tests pass, with 7 optional-integration skips on Linux/CPython 3.12.git diff --check: clean.Draft until the updated branch's complete GitHub Actions matrix passes. Native matching and scientific policies are unchanged.