Skip to content

Latest commit

 

History

History
73 lines (50 loc) · 11.4 KB

File metadata and controls

73 lines (50 loc) · 11.4 KB

2026-09 update log (e)

Index and query commands: README.md. New entries go at the end.


U-20260926-35 · 2026-09-26 · Buffer an unterminated SSE line in pieces instead of re-splitting the whole buffer on every chunk · #bugfix #performance #http

Conformance run of utils/sse_client against the WHATWG event-stream algorithm and the WPT eventsource/format-* streams: parsing was already correct on all 19 vectors and on 30,000 randomly chunked streams. Its cost was not.

  • Quadratic feed: SSEParser.feed joined the buffered partial line with each new chunk and re-split the whole of it, so one long data line cost the square of its length: a 2 MB line fed in 1 KB chunks took 10.75 s (1 MB: 2.0 s). A chunk with no line break is now appended to a list of pieces and nothing else happens; the pieces are joined once, when a break arrives. Measured: 2 MB in 0.02 s, 1 MB in 0.01 s. The WPT vectors and the chunking fuzz still match exactly.
  • test_sse_client_chunking.py (new, 3) fails the 2 MB timing case on the old code; the 75 existing SSE and HTTP tests pass.
  • Not changed: close() and parse_event_stream still dispatch an event the stream ended without a blank line for. WHATWG discards it and WPT format-data-before-final-empty-line expects that, but the flush is documented behaviour that existing callers read complete blobs with.
  • Files: utils/sse_client/sse_client.py, the new test, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-36 · 2026-09-26 · Validate a multipart boundary, refuse or redraw one a part contains so a field value cannot inject a part, and read only the boundary parameter · #bugfix #security #http

Checked utils/multipart against RFC 2046 5.1.1 and RFC 7578. test_multipart_rfc2046.py (new, 14) fails 10 on the old code; the 50 existing multipart and HTTP tests pass.

  • Part injection: nothing checked that a part's content was free of the delimiter. build_multipart({"a": 'x\r\n--B\r\nContent-Disposition: form-data; name="evil"...'}, boundary="B") parsed back as two fields, one named by the value. RFC 2046 5.1.1: the delimiter must not appear in an encapsulated part. A caller's boundary that a field value or file content contains raises MultipartError; a generated one is redrawn until no part contains it.
  • Header injection: the boundary was used as given, so boundary="B\r\nX-Injected: 1" wrote a second header line into the returned Content-Type. It must now be 1-70 of letters, digits and '()+_,-./:=? or space, not ending in a space, and one that is not a token is quoted in the Content-Type.
  • Wrong parameter: _boundary_of matched boundary= anywhere, so notboundary=x; boundary=y read x. It is anchored to a parameter start.
  • MultipartError is an AutoControlException and a ValueError, so existing except ValueError callers keep working; the content-type line-break check and the missing-boundary error raise it too.
  • Files: utils/multipart/{multipart,__init__}.py, the new test, docs/source/{Eng,Zh}/doc/new_features/v89_features_doc.rst, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-37 · 2026-09-26 · Leave untranslated entries out of a compiled .mo, as msgfmt does, so gettext readers fall back to the msgid instead of showing an empty string · #bugfix #i18n

Differential test of utils/gettext_catalog over 400 random catalogs (contexts, plurals, split strings, escapes, fuzzy flags): our .po parse against babel's read_po, our .mo read by Python's gettext.GNUTranslations, and babel's .mo read by our read_mo. Everything agreed except one class.

  • Empty strings in the UI: to_mo_bytes / compile_mo wrote an entry whose msgstr was empty with an empty translation. The catalog's own gettext falls back to the msgid, but every .mo reader, Python's gettext included, returned "" for it, so an untranslated label showed as blank. GNU msgfmt leaves such entries out; _mo_pairs now skips an entry whose first form is empty (fuzzy ones were already skipped at parse time). The differential test goes from 107 mismatches to 0.
  • test_gettext_catalog_msgfmt.py (new, 2) fails 1 on the old code; the 50 existing gettext tests pass.
  • Files: utils/gettext_catalog/gettext_catalog.py, the new test, docs/source/{Eng,Zh}/doc/new_features/v114_features_doc.rst, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-39 · 2026-09-26 · Count a yearly BYDAY ordinal within the year alongside BYMONTHDAY, keep DTSTART's microseconds, let count/until narrow a rule, apply DAILY BYSETPOS, end a series by its gap instead of a fixed 400 years, and refuse the rule parts RFC 5545 forbids · #bugfix #scheduling

Three-way run of utils/recurrence over 6,068 rules (68 targeted, including the RFC 5545 3.8.5.3 examples, plus 6,000 random): the module, python-dateutil 2.9, and a brute-force reference written from the RFC. Where dateutil and the RFC disagree (529 mixed plain/numbered BYDAY lists, 17 WEEKLY+BYSETPOS) the module already followed the RFC. After this change it agrees with the reference on all 6,068; before, 26 differed. test_recurrence_rfc5545.py (new, 19) fails 17 on the old code; the 52 existing recurrence tests pass.

  • Year-relative ordinals: FREQ=YEARLY;BYMONTHDAY=1;BYDAY=1MO without BYMONTH counted 1MO within each month (2023-05-01, 2024-01-01, ...). Per 3.3.10 the offset is within the year when BYMONTH is absent: 2024-01-01, 2029-01-01, 2035-01-01, 2046-01-01.
  • Microseconds: _apply_time dropped them, so a DTSTART of 09:00:00.000001 lost its first occurrence, which the RFC says DTSTART always is.
  • count / until: they replaced the rule's own COUNT / UNTIL, and AC_rrule_occurrences / ac_rrule_occurrences pass count=10, so FREQ=DAILY;COUNT=3 listed ten dates. The smaller limit wins now.
  • DAILY BYSETPOS was ignored (BYMONTH=1;BYSETPOS=2 gave every January day); it applies to the one-day set.
  • Silent truncation: the series stopped 400 years after DTSTART and after 100,000 candidates. YEARLY;INTERVAL=100;COUNT=3 from 2000-02-29 lost 2800, COUNT=150000 gave 100,000, and next_occurrence of a daily rule from 1750 answered None. A series now ends at 9999-12-31 or after 400 years without an occurrence (scaled for periods longer than a year); max_iter caps only rules without a count, and next_occurrence walks to now.
  • Validation: BYMONTHDAY with WEEKLY (a MUST NOT), a numbered BYDAY with DAILY / WEEKLY, and WKST=XX (which became MO) raise AutoControlException.
  • Errors outside the family: UNTIL=2024-01-03 raised ValueError, stepping past year 9999 raised OverflowError, and a naive now against an aware DTSTART raised TypeError. The first is a rule error now, stepping stops at the end of the calendar, and now is read the way UNTIL is.
  • Files: utils/recurrence/recurrence.py, the new test, docs/source/{Eng,Zh}/doc/new_features/v66_features_doc.rst, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-38 · 2026-09-26 · Encode a bytes attribute as OTLP bytesValue instead of its Python repr, and recognise problem+json only as the Content-Type media type · #bugfix #observability #http

Two small spec gaps found while checking the observability and HTTP helpers against their standards. test_otlp_problem_media.py (new, 8) fails 4 on the old code; the 64 existing OTLP, problem and HTTP tests pass.

  • OTLP: the output of spans_to_otlp parses cleanly with the official opentelemetry-proto definitions (json_format.Parse, unknown fields refused): int64 as strings, enums as integers, arrays and kvlists nested. A bytes attribute, though, went through str() and landed as stringValue "b'\\x00\\x01'". OTLP AnyValue has bytes_value, base64 in the protobuf JSON mapping; bytes / bytearray are written that way now.
  • RFC 9457: is_problem searched the Content-Type for application/problem+json, so text/plain; note=application/problem+json and application/problem+json-seq were treated as problem documents and parsed. It compares the media type before ; now (case-insensitive).
  • Files: utils/otlp_export/otlp_export.py, utils/http_problem/http_problem.py, the new test, docs/source/{Eng,Zh}/doc/new_features/{v78,v94}_features_doc.rst, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-40 · 2026-09-26 · Give French its CLDR ordinals, match a locale by language and fall back to Babel instead of English, compare =N selectors as numbers, and reject the patterns ICU rejects · #bugfix #i18n

Checked utils/message_format against CLDR (through Babel) and the ICU MessageFormat syntax. English cardinal and ordinal and French cardinal categories already matched for integers, decimals and negatives, and apostrophe quoting, nested # and offset: were right. test_message_format_icu.py (new, 26) fails 25 on the old code; the 61 existing message-format and text tests pass.

  • French ordinals used the English rules, which the docs say locale="fr" selects: 21 rendered 21er and ordinal_category(2, "fr") was two. CLDR's French ordinal is one for n = 1 only.
  • Locales: the lookup was an exact key, so fr_FR, fr-CA and FR got English rules (0 chats, and a million was other, not many), and every other locale silently got English (plural_category(2, "ru") was other; CLDR says few). The language subtag selects the built-ins now; other locales use Babel's CLDR data when Babel is installed and raise MessageFormatError when it is not.
  • Syntax ICU rejects was accepted and rendered something: {n, plural, one {x} rendered x, {n, plural} tail {x} swallowed the tail, one x other {y} rendered " other ", a message with no other rendered "", and a late offset: or a duplicate selector passed. Each raises MessageFormatError now.
  • =N selectors compared as text, so =1.0 never matched 1; they compare as numbers.
  • Errors: parse and render failures raised ValueError and plural_category(float("inf")) raised OverflowError, outside the AutoControlException family. MessageFormatError is both, and an infinite value is other (rendering it as a count is still an error).
  • Files: utils/message_format/{message_format,__init__}.py, the new test, docs/source/{Eng,Zh}/doc/new_features/v113_features_doc.rst, CHANGELOG.md, architecture_explore.md (line counts).

U-20260926-41 · 2026-09-26 · Pin ruff's rule set in pyproject.toml so a ruff upgrade does not change what CI enforces · #ci #tooling

Dependabot's PR #476 (ruff 0.15.22 to 0.16.0) has stayed open: ruff 0.16 widened its default rule set to pyupgrade, isort, RUF, SIM, DTZ, TRY, BLE and more, and the same tree reported 13,934 findings under it (6,312 UP006, 3,419 UP045, 1,506 I001, 1,397 UP035, 611 RUF100, ...) against none under 0.15. [tool.ruff.lint] only extended the default with E501, so what CI enforced depended on the ruff version.

  • select = ["E4", "E7", "E9", "F", "E501"] names the rule set explicitly: ruff's pre-0.16 default, which this code was linted against, plus the documented line limit. ruff check je_auto_control/ passes under both 0.15.22 (CI's pin) and 0.16.0, with identical results on je_auto_control/ and test/.
  • Adopting the newer rules is left to the maintainer; mass-rewriting 13k annotations would touch modules a running Jeffrey_RPA batch imports.
  • CLAUDE.md said CI runs ruff "with its default rules"; it names the pinned set now.
  • Files: pyproject.toml, CLAUDE.md.