Skip to content

Backport bug fixes to 3.x for v3.38.0 - #1104

Merged
laughingman7743 merged 20 commits into
3.xfrom
backport/3.x-v3.38.0
Oct 6, 2026
Merged

laughingman7743 merged 20 commits into
3.xfrom
backport/3.x-v3.38.0

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

WHAT

This PR backports bug fixes from master to the 3.x maintenance branch for v3.38.0. It includes the fixes for the externally reported issues #854 and #857.

Each master PR is cherry-picked with -x in master merge order, one commit per PR. When a conflict came from a master PR that is not backported, only the backported PR's own changes are applied. Each commit message records how its conflict was resolved.

PR Fix 3.x handling
#800 to_sql finds an existing table whatever the case of its name (#799) clean
#817 Cache-hit tests use ENV.work_group instead of primary (#816) clean
#831 Spark session readiness polling stops on TERMINATED, DEGRADED, or FAILED (#792) docstring context from #803
#830 AsyncSparkCursor.close() shuts down its executor when session termination fails (#795) clean
#832 A Spark session that a cursor starts is terminated when cursor setup fails (#794) docstring and annotation context from #803
#899 Synchronous iteration of the asyncio cursors and result sets raises TypeError instead of looping forever (#898) applied to 3.x's pre-#883 WithAsyncFetch
#900 Arrow decimal256 maps to Athena decimal instead of string (#888) decimal256 only (see below)
#928 Stale S3FileSystem listing and object caches (#922) clean
#929 Appending to an object keeps its data; this fixes data loss, truncation by an empty append, and duplicated bytes (#921) docstring context from #919
#947 A single large write() no longer leaves a sub-5 MiB non-last part, which made the upload fail with EntityTooSmall (#942) clean
#955 Multipart copy parts stay within the S3 part size limits (#951) test import from #948
#959 The client-side result cache is skipped for qmark queries with parameters (#941) applied to 3.x's _execute()
#950 as_pandas() returns the whole result when auto_optimize_chunksize chose a chunk size (#924) docstring and docs context from #919, #940
#995 Top-level MAP and STRUCT/ROW columns reflect as AthenaMap / AthenaStruct (#857) visit_null kept (see below)
#1023 ArrowCursor fallback values are not converted twice (#1021) clean
#1029 A closed PolarsDataFrameIterator stops for every reader kind (#1019) test file created on 3.x
#1031 ArrowCursor keeps the NULL rows of single-column CSV results (#1026) clean
#1033 A numeric UTC offset in TIMESTAMP WITH TIME ZONE gives an aware datetime (#854) #854 part only (see below)
#1045 ArrowCursor reads multi-line CSV values that cross a read block (#1040) clean
#1075 PandasCursor uses the C engine for DDL .txt results with engine="pyarrow" (#1073) docs context from #902 / #1057, test context from #1001

Ported in part, by maintainer decision:

Not backported:

WHY

master is 4.0.0 development. Many bugs fixed there since v3.36.0 also exist on 3.x.

v3.38.0 ships the fixes that are safe for the 3.x line, without its breaking changes. It includes the fixes for the issues @aminghadersohi reported: #854 (TIMESTAMP WITH TIME ZONE with a numeric offset came back naive) and #857 (top-level MAP columns reflected as String).

Release notes for v3.38.0 (in addition to the unreleased #839, #905, #906, #913 already on 3.x):

TEST

All runs were at 5bf0490. The branch head 683658f has the same tree; only three commit messages changed afterwards.

  • just lint (ruff check, ruff format --check, mypy): passed.
  • markdownlint-cli2 on the changed docs: 0 errors.
  • Offline, on Python 3.10.16 with the 3.x lock (pandas 2.2.3), using --noconftest: 342 passed. The run covered:
    • tests/pyathena/test_converter.py
    • aio/test_common.py and aio/test_result_set.py
    • spark/test_common.py
    • arrow/test_util.py
    • pandas/test_result_set.py and polars/test_result_set.py
    • sqlalchemy/test_compiler.py and sqlalchemy/test_base.py::TestAthenaDialect
  • Real AWS, in the maintainer's test account, Python 3.10.16, -n 8: 1178 passed.
    • The run covered tests/pyathena/test_cursor.py and aio/test_cursor.py, tests/pyathena/pandas/, arrow/, polars/, filesystem/, sqlalchemy/test_base.py, test_converter.py, the aio result-set and common tests, and spark/test_common.py.
    • On pandas 2.2.3, PandasDataFrameIterator.as_pandas() (Return the whole result from as_pandas() when auto chunking applies #950) emits pandas' FutureWarning about concatenating all-NA chunks in test_as_pandas_matches_whole_read[category_null_chunk]. Master does not show it because it requires pandas 3; the test passes.
  • Not run locally: the Spark integration suite, the full tests/pyathena/ suite, and just test sqla / sqla-async. The PR's AWS jobs run them on the newest Python once the PR is Ready, and the release gate runs every Python version for the tag.

🤖 Generated with Claude Code

laughingman7743 and others added 11 commits October 6, 2026 21:03
(cherry picked from commit ad1232b)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imary

(cherry picked from commit 0c26c7e)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 37999c3)

Conflict resolution: pyathena/spark/common.py conflicted in the
_exists_session() docstring, which master gained with the metadata
work of #803 (not backported to 3.x). #831 only extended that
docstring's Raises entry, so 3.x keeps its existing undocumented
method there. The _wait_for_idle_session() fix and its new docstring
apply unchanged, as do the tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ermination fails

(cherry picked from commit 2a107ad)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tup fails

(cherry picked from commit 5766a13)

Conflict resolution: pyathena/spark/common.py conflicted in
_terminate_session(), whose docstring and dict[str, Any] request
annotation master gained with the metadata work of #803 (not
backported to 3.x). Only #832's own changes are applied there:
_terminate_session() delegates to the new __terminate_session(), which
takes the session ID, keeps 3.x's unannotated request dict, and logs
the session ID. The __init__ reordering, the _start_session() cleanup,
the AsyncSparkCursor executor setup, and the tests apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d result sets

(cherry picked from commit 5415c05)

Conflict resolution: pyathena/aio/common.py conflicted because 3.x
still has the pre-#883 WithAsyncFetch, which also subclasses
CursorIterator and carries default synchronous fetch methods, and its
imports and docstring differ. Only #899's own changes are applied:
NoReturn is added to 3.x's typing import, the class docstring gains
"Synchronous iteration raises ``TypeError``.", and __iter__() is added
after the synchronous fetch methods. Every 3.x asyncio SQL cursor
overrides the fetch methods with coroutines, so a synchronous for loop
hung there too. pyathena/aio/result_set.py and the tests apply
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 8679f04)

3.x resolution: only the decimal256 part of #900 is applied. 3.x
allows pyarrow>=10.0.0, but pyarrow.types.Type_DECIMAL32 and
Type_DECIMAL64 first appear in pyarrow 19.0.0; referencing them would
raise AttributeError in get_athena_type() for every type that reaches
the decimal check on older pyarrow. The check therefore lists
Type_DECIMAL128 and Type_DECIMAL256 (replacing the Decimal256Type class
that never matched a type id), and the decimal32/decimal64 test cases
are dropped. The decimal128 and decimal256 cases are kept unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 7be04bd)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r block size

(cherry picked from commit f8e6b86)

Conflict resolution: pyathena/filesystem/s3.py conflicted in the
S3File.__init__() docstring, which master gained with #919 (not
backported to 3.x). #929 only reworded that docstring's append-mode
sentence, so 3.x keeps its undocumented __init__(). The append, upload,
and discard fixes and the tests apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rt part

(cherry picked from commit e0e85da)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c9d6875)

Conflict resolution: tests/pyathena/filesystem/test_s3_async.py
conflicted in its imports. Master already imported unittest.mock there
from #948 (not backported to 3.x), next to which #955 added
SimpleNamespace. #955's new test uses both, so both imports are added.
The runtime changes and the other tests apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread pyathena/converter.py
_UTC_OFFSET_PATTERN: re.Pattern[str] = re.compile(r"([+-])(\d{2}):(\d{2})")


def _parse_utc_offset(value: str) -> timezone | None:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 1: implementation behavior: CLEAN (no actionable findings)

Scope: git diff f09cf4269052b7dfa4157f10a5f37a5cbfa6148a..5bf0490e450e56ca9a1be3840e94d86234beae4f, the 20 commits, 36 files.

Covered:

Recorded limitations (3.x consequences of the agreed scope, not defects of this PR) are in the threads below.

Comment thread pyathena/aio/common.py
result_set = cast(AthenaResultSet, self.result_set)
return result_set.fetchall()

def __iter__(self) -> NoReturn:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 note (#899):

except (TypeError, ValueError):
util.warn(f"Did not recognize type '{type_}'")
return types.NullType()
if not _nested and name in ("map", "row", "struct") and length:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 note (#995):

  • visit_null is kept, so NullType is not rejected in DDL on 3.x.
  • A reflected MAP or STRUCT with an unrecognized nested type renders NULL in CREATE TABLE, and Athena rejects the statement. Master raises CompileError at compile time instead.
  • Before this change, the whole top-level column reflected as String.
  • 3.x's AthenaArray(NullType) already behaved this way.
  • This is the agreed partial port: removing visit_null without Split SQLAlchemy type rendering into Hive DDL and Trino DML compilers #961 would break CAST(x AS NULL).

laughingman7743 and others added 9 commits October 6, 2026 21:32
(cherry picked from commit f3ef301)

Conflict resolution: 3.x's BaseCursor._execute() and
AioBaseCursor._execute() still build the request inline, without the
docstring and _build_execute_request() that master gained with #883
and the _start_execution() from #853 (neither backported to 3.x). #959's
own change is applied to 3.x's code: the _find_previous_query_id()
lookup runs only when the request has no ExecutionParameters, with the
same comment. The docstring sentence it added to _execute() has no 3.x
docstring to go into; the ExecuteOptions docstring line applies
unchanged. docs/usage.md gains the same sentence after 3.x's (older)
cache paragraph. In tests/pyathena/aio/test_cursor.py only #959's
test_execute_qmark_parameters_skip_cache is added, without the #853
interrupt tests around it. In tests/pyathena/test_cursor.py,
test_cache_size_with_qmark_parameters uses datetime.now(timezone.utc)
as the rest of 3.x's file does, because master's datetime.UTC comes
from #884 and needs Python 3.11; the tests otherwise apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nking applies

(cherry picked from commit 700fa30)

Conflict resolution: pyathena/pandas/result_set.py conflicted in the
AthenaPandasResultSet.as_pandas() docstring, which master gained with
#919 (not backported to 3.x); 3.x keeps its undocumented method, and
the code change and the other docstring updates apply unchanged.
docs/aio.md conflicted because master's paragraph on the synchronous
convenience methods was rewritten by #940 (not backported); #950's
sentence is added after 3.x's paragraph. docs/pandas.md and the tests
apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…and MAP columns

(cherry picked from commit 6be7de3)

3.x resolution: the reflection fix for #857 is applied unchanged
(pyathena/sqlalchemy/base.py and its docstring). The part of #995 that
removed AthenaTypeCompiler.visit_null, so that NullType raises
CompileError in DDL, is left out by maintainer decision. It relies on
the DDL/DML type compiler split of #961 (master only): 3.x still
renders CAST types through AthenaTypeCompiler, so without visit_null
CAST(x AS NULL) would raise. pyathena/sqlalchemy/compiler.py is
therefore unchanged, and so are the DDL-rejection documentation and
tests:
- docs/sqlalchemy.md: the "DDL and CAST types" sentence on NullType is
  dropped together with the #961 section it belongs to; the paragraphs
  on reflected STRUCT/ROW and MAP columns apply unchanged.
- tests/pyathena/sqlalchemy/test_compiler.py: only
  test_null_type_cast_is_unchanged is added, after
  test_ddl_and_cast_types as on master; test_null_type_is_rejected_in_ddl
  is dropped.
- tests/pyathena/sqlalchemy/test_base.py: the new dialect tests go into
  3.x's TestAthenaDialect (from #839) after its existing tests, without
  the master-only metadata tests around them, and need the AthenaDialect
  import. test_unrecognized_field_type_reflects_null_type_and_blocks_ddl
  keeps its reflection assertions without the CompileError check and is
  named test_unrecognized_field_type_reflects_null_type. The added
  assertion in test_columns_from_information_schema is dropped, because
  that test comes from #777 (not backported). The TestSQLAlchemyAthena
  changes apply unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 52ea5a7)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r kind

(cherry picked from commit 9a373fd)

Conflict resolution: tests/pyathena/polars/test_result_set.py does not
exist on 3.x; master created it with #828 (not backported). The file is
created with master's header, the imports #1029's test needs, and
#1029's TestPolarsDataFrameIterator; #828's chunk-reader tests are not
included. pyathena/polars/result_set.py applies unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ursor

(cherry picked from commit 3323784)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…H TIME ZONE values

(cherry picked from commit ec5323e)

3.x resolution: only the #854 fix is ported from #1033, by maintainer
decision: a TIMESTAMP WITH TIME ZONE value with a numeric UTC offset
(+05:30, -08:00) converts to an aware datetime with a fixed-offset
time zone instead of a naive one. #1033's other changes are behavior
changes for 4.0.0 and build on #1005 and #1010 (not backported), so
they are left out: TIME WITH TIME ZONE conversion, empty TIME, JSON and
TIMESTAMP WITH TIME ZONE text as NULL, the time zone converters of the
pandas, Arrow and Polars cursors, the type-signature and JSON element
changes in parser.py, the docs, and the live CONVERTED_VALUES tests.

Applied from #1033:
- pyathena/converter.py: _UTC_OFFSET_PATTERN and _parse_utc_offset()
  as on master, and _to_datetime_with_tz() uses
  _parse_utc_offset(tz) or gettz(tz). It keeps 3.x's strptime() parsing
  (master parses with _parse_datetime() from #818, not backported) and
  its None check.
- tests/pyathena/test_converter.py:
  test_to_datetime_with_tz_offsets_and_zone_names with the imports it
  needs. The "" case (empty text as NULL) and the case without
  fractional seconds (needs #818's parser) are dropped; Athena renders
  these values with milliseconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…wCursor

(cherry picked from commit 50b8718)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit bf6f299)

Conflict resolution: the _get_csv_engine() fix in
pyathena/pandas/result_set.py applies unchanged. The other files
conflicted with text from PRs that are not backported to 3.x:
- docs/pandas.md: master's paragraph on when engine="pyarrow" applies
  was added by #902 and extended by #1033 and #1057; 3.x has none of
  that text. #1075 added a ".txt" condition to its first sentence,
  which has no 3.x counterpart, and a sentence on DDL results; only
  that sentence is added, after the read_csv options.
- tests/pyathena/pandas/test_cursor.py: only #1075's
  result_set._query_execution = None is added; the _metadata and
  _kwargs lines around it come from other PRs.
- tests/pyathena/pandas/test_result_set.py: 3.x's file (from #950) has
  no TestAthenaPandasResultSet, so the class is created with #1075's
  test_get_csv_engine_result_format and the imports it needs.
  test_read_csv_ddl_preserves_numeric_looking_names is dropped: it
  stubs result_set._fs.open() and asserts that the stream is closed,
  but 3.x's _read_csv() passes the S3 path to pandas.read_csv(), and
  reading through the result set's filesystem comes from #1001 (not
  backported).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread pyathena/arrow/util.py
if type_.id == types.Type_TIMESTAMP: # 18
return "timestamp", 3, 0
if type_.id in [types.Type_DECIMAL128, types.Decimal256Type]: # 23, 24
if type_.id in [types.Type_DECIMAL128, types.Type_DECIMAL256]: # 23, 24

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 2: claims, callers and operations: FINDINGS, repaired (wording in commit messages and the PR body only; no code change)

Scope: git diff f09cf4269052b7dfa4157f10a5f37a5cbfa6148a..683658f24dff765e0cc6df1bffcfebdd9f7797a8, the PR body, and the 20 commit messages.

Claims checked against evidence:

Comment thread docs/aio.md

The `as_pandas()`, `as_arrow()`, and `as_polars()` convenience methods operate on
already-loaded data and remain synchronous.
With a chunk size chosen by `auto_optimize_chunksize`, `as_pandas()` reads every remaining chunk.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 deferral (docs reader):

for row in zip(*dict_rows.values(), strict=False)
]
else:
processed_rows = list(zip(*dict_rows.values(), strict=False))

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): Codex CLI 0.160.0, model gpt-6.1-sol, codex exec -s read-only, effort high. Session 01a11134-c83c-7aa2-96a0-26d06ed51cb0. Static review only: no builds, tests, edits or network.

  • Snapshot: detached worktree at head 683658f24dff765e0cc6df1bffcfebdd9f7797a8, base f09cf4269052b7dfa4157f10a5f37a5cbfa6148a. The prompt contained the diff range, the mapping from commit to master merge commit, the 3.x dependency floors and the required scope. It contained no PR text, commit messages or prior findings.
  • The snapshot and the PR worktree were unchanged afterwards.
  • Coverage reported:
    • all 20 commits against their master patches, including the three partial ports;
    • sync, threaded and asyncio cursors; pandas, Arrow, Polars and S3FS result sets;
    • Spark lifecycle;
    • S3 caches, append, multipart upload/copy and failure paths;
    • SQLAlchemy reflection and compiler, dependency bounds, tests and docs.
  • Verdict: FINDINGS, with two P2 findings: one introduced (this thread) and one pre-existing (next thread).

Finding (introduced, P2):

  • With managed result storage (no output location), the GetQueryResults fallback no longer applies the result set's converter mapping in the fetch methods.
  • So a custom mapping such as DefaultArrowTypeConverter().set("varchar", str.upper) stops applying: fetchall() returns ("hello",) instead of ("HELLO",). The same holds on master since Do not convert ArrowCursor fallback values twice #1023.

Author verification: confirmed, deferred.

self._sync_fs.max_workers,
block_size,
)
ranges = self._sync_fs._get_copy_ranges(size1, block_size)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding (pre-existing, P2), relayed from the independent review:

  • Location: pyathena/filesystem/s3_async.py:276, AioS3FileSystem._copy_object_with_multipart_upload.
  • After CreateMultipartUpload, an UploadPartCopy error, a cancellation, or a completion failure propagates without AbortMultipartUpload, so the uploaded parts are left behind.

Author verification: pre-existing at the review base, not introduced.

@laughingman7743
laughingman7743 marked this pull request as ready for review October 6, 2026 12:51
def as_pandas(self) -> PandasDataFrameIterator | DataFrame:
if self._chunksize is None:
return next(self._df_iter)
return self._df_iter.as_pandas()

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Second independent review (relayed, requested by the maintainer): Claude Code CLI, model claude-fable-5-1, max profile (first-party Max auth verified), effort high, --permission-mode plan.

  • Tools: read and git only (Read/Grep/Glob, git show/diff/log/grep/blame). Session 40ff7b56-3d85-4a99-88a5-7cebd4deb23d.
  • Static review on a detached snapshot at 683658f24dff765e0cc6df1bffcfebdd9f7797a8 (base f09cf4269052b7dfa4157f10a5f37a5cbfa6148a). The snapshot was unchanged afterwards.
  • The prompt matched the Codex one, plus a compatibility audit:
    • breaking changes;
    • every user-visible behavior change, classified (a) fix of a wrong or failing result, or (b) change to previously correct, working behavior.

Coverage reported:

  • all 20 commits against master, hunk by hunk;
  • cursor bases and the cache;
  • aio iteration, including pyathena/aio/sqlalchemy, which awaits fetchall() and never iterates;
  • Spark lifecycle;
  • the Arrow, pandas and Polars result sets;
  • the converters, with gettz("+05:30") is None confirmed on dateutil 2.9;
  • SQLAlchemy reflection traces;
  • the filesystem cache, append, multipart and copy paths, against the fsspec 2024.12 flush() contract;
  • the tests (each traced to fail without its fix);
  • the docs.

Result:

  • Correctness, faithfulness, tests and docs: CLEAN. No defects introduced.
  • Breaking changes: none. No public name removed or renamed; no signature, default, return shape or dependency floor changed. New exceptions occur only on paths that used to hang, loop forever or return wrong data.

Category (b) behavior changes it listed:

Pre-existing 3.x defects it noted (not introduced; fixed on master by excluded PRs):

Minor docs note: docs/aio.md:192-194, already recorded in round 2.

Author verification and actions:

  • The Return the whole result from as_pandas() when auto chunking applies #950 second-call behavior was checked against the code: 3.x used next(self._df_iter), and PandasDataFrameIterator.as_pandas() returns pd.DataFrame() when nothing remains.
  • All category (b) items were added to the release notes in the PR body.
  • No code change: each item is either the documented contract of its fix or the same behavior as master.

@laughingman7743
laughingman7743 merged commit 7fc7f07 into 3.x Oct 6, 2026
12 checks passed
@laughingman7743
laughingman7743 deleted the backport/3.x-v3.38.0 branch October 6, 2026 15:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant