Skip to content

perf: accelerate exact count-table parsing and preserve strict input boundaries - #149

Merged
dnncha merged 3 commits into
mainfrom
perf/python-input-throughput-20260930
Oct 4, 2026
Merged

dnncha merged 3 commits into
mainfrom
perf/python-input-throughput-20260930

Conversation

@dnncha

@dnncha dnncha commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Ordinary raw-count cells previously passed through regex and Decimal parsing. Parse unsigned ASCII integers of at most 64 characters directly as exact integers, preserving counts above 2**53. Longer inputs, signed values, decimals and scientific notation retain the existing Decimal path and limits; Unicode digits remain invalid and unselected sample columns remain validated.

The production diff against current main is the five-line count-parser fast path. Main already contains the equivalent FASTQ printable-ASCII optimization, and this branch retains main's implementation. Added 537 boundary cases cover byte-range characters, Unicode, lossless counts, conversion limits and unselected columns.

The reproducible benchmark and raw evidence in docs/audits/2026-09-30-input-throughput.{md,json} record the original 30 September baseline/candidate pair. Five alternating Linux/CPython 3.12 runs measured count-table parsing at 0.468 → 0.176 seconds (2.66x throughput) for 10,000 guides × 32 samples, with identical complete parsed-output hashes. These are synthetic Python parsing measurements, including output hashing; they do not establish whole-assay speedups. The historical FASTQ measurement describes the optimization already integrated into main. These measurements were not rerun during reconciliation.

Current integration validation incorporates main through 2fa3532549d5945991aba2d7645af66108206497, including the strict native gzip fix:

  • make all shared test cli-test python-test: native build, both C test executables, CLI fixtures and native/Python FASTQ parity pass; 2,069 Python tests pass, with 7 optional-integration skips on Linux/CPython 3.12.
  • git diff --check: clean.
  • The original candidate also passed reproducible source-distribution/wheel and installed CLI lifecycle checks; all 537 added boundary cases passed against the original baseline functions.

Draft until the updated branch's complete GitHub Actions matrix passes. Native matching and scientific policies are unchanged.

@dnncha dnncha changed the title perf: accelerate strict Python FASTQ and count-table parsing perf: accelerate exact count-table parsing and preserve strict input boundaries Oct 3, 2026
@dnncha
dnncha marked this pull request as ready for review October 4, 2026 08:45
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-04T08:47:49.401860Z ba88263 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@dnncha
dnncha merged commit b7be660 into main Oct 4, 2026
35 of 37 checks passed
@dnncha
dnncha deleted the perf/python-input-throughput-20260930 branch October 4, 2026 08:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant