Skip to content

Require pandas 3.0 and pyarrow 22.0 - #885

Draft
laughingman7743 wants to merge 2 commits into
feat/864-drop-python-310from
feat/852-pandas-3-pyarrow-22
Draft

laughingman7743 wants to merge 2 commits into
feat/864-drop-python-310from
feat/852-pandas-3-pyarrow-22

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

WHAT

Raise the pandas and Arrow extras to versions the project tests, for 4.0.0. Stacked on #884 (drop Python 3.10); the base will be retargeted to master once #884 merges.

  • pandas extra: pandas>=3.0.0 (was >=1.3.0 below Python 3.13, >=2.3.0 from 3.13).
  • arrow extra: pyarrow>=22.0.0 for every Python version (was >=10.0.0 below Python 3.14).
  • The dev group follows the extras and drops its numpy entries. pandas 3.0 requires numpy>=1.26.0 (>=2.3.3 on Python 3.14) itself, and the removed markers only restated that. They also left Python 3.13 unconstrained. The tests still import numpy, which pandas installs.
  • README.md and docs/introduction.md list the new minimums. docs/arrow.md drops a note saying the ArrowCursor S3 timeouts need PyArrow 10.0.0 or later, since every supported version has them.
  • docs/null_handling.md and the pandas cursor tests (tests/pyathena/pandas/test_cursor.py, test_async_cursor.py) drop their pandas 2 branches. String columns are the pandas str dtype, and NULL is NaN. The future.infer_string opt-out stays documented because pandas 3.0.0 and 3.0.6 still accept it.
  • uv.lock: only the declared specifiers change; the resolved versions do not.

The library itself has no pandas or pyarrow version branches to remove. Its ImportError handling only detects whether an optional package is installed.

Release note (breaking): the pandas extra requires pandas 3.0.0 or later, and the arrow extra requires pyarrow 22.0.0 or later. Users who need pandas 2 or an older pyarrow can stay on the 3.x maintenance branch.

WHY

Closes #852. The declared floors were never installed by CI, so they promised compatibility that was not verified. In #828, an untested floor hid a real defect: with Polars 1.34–1.36, a chunked read that failed partway ended without an error (#820).

  • pandas: per the decision on Raise the pandas and pyarrow minimum versions to tested versions #852, 4.0.0 supports pandas 3.x only. pandas 3.0 requires Python 3.11, which Drop Python 3.10 support #884 makes the minimum.
  • pyarrow: 22.0.0 (2025-10-24) is the first release with wheels for every supported Python version, 3.11–3.14. 21.0.0 has no cp314 wheels. One floor without markers can therefore be verified on every supported version. It is also above the pyarrow 13.0.0 that pandas 3.0 requires for its pyarrow features.

TEST

Tested commit: 29bfde0

  • just format, just lint: pass. just docs lint: pass. just docs build: succeeds, with no warnings from the changed pages.
  • Floor verification against AWS: separate venvs for Python 3.11.11 and 3.14.5, each with the project, its extras, and the dev group, and with exactly pandas 3.0.0, pyarrow 22.0.0, and polars 1.39.0 (numpy 2.4.6 / 2.5.3). Versions were checked before each run. Command: UV_PROJECT_ENVIRONMENT=<venv> uv run --no-sync --env-file .env python -m pytest -n 8 --reruns 1 --only-rerun … tests/pyathena/{pandas,arrow,polars} tests/pyathena/aio/{pandas,arrow,polars} tests/pyathena/sqlalchemy/test_base.py::TestSQLAlchemyAthena::{test_to_sql_parquet,test_to_sql_json,test_to_sql_column_options}.
    • Python 3.11: 485 passed. Python 3.14: 485 passed. Both runs raised the same 2 ParserWarnings from test_binary_custom_dialect, which passes a CSV dialect whose quoting overrides PyAthena's.
  • Locked versions (pandas 3.0.6, pyarrow 25.0.1): covered by the PR's AWS run once Ready.

Not covered: numpy at pandas' own floor (1.26.0), and the fastparquet engine, which is not a dependency of any extra.

🤖 Generated with Claude Code

laughingman7743 and others added 2 commits September 29, 2026 01:38
The pandas and arrow extras declared floors that CI never installed
(pandas>=1.3.0, pyarrow>=10.0.0), so they promised compatibility that
was not verified. With Python 3.10 dropped, the extras can require the
versions the project tests:

- pandas>=3.0.0, the first release of the only major version CI tests.
- pyarrow>=22.0.0, the first release with wheels for every supported
  Python version, including 3.14.

The dev group follows the extras and no longer lists numpy, whose
markers only restated pandas' own numpy requirement. The pandas cursor
tests and the NULL handling guide drop their pandas 2 branches.

Closes #852

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every supported pyarrow version accepts the S3FileSystem timeouts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread pyproject.toml
pandas = [
"pandas>=1.3.0; python_version<'3.13'",
"pandas>=2.3.0; python_version>='3.13'",
"pandas>=3.0.0",

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): FINDINGS (1, repaired)

Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 (the head of #884; this PR is stacked on it) .. head 29bfde0722f3674883960899e0eb84dc80bb7378, all 8 changed files.

Finding:

  • docs/arrow.md:296-298 (as of 2be0c42) still said the ArrowCursor connect_timeout/request_timeout options "require PyArrow >= 10.0.0". With the floor at 22.0.0, every installable version has them, so the note was obsolete. It was removed in 29bfde0. A git grep for pandas/pyarrow/numpy version mentions outside the lock finds nothing else.

Covered, no finding:

  • The library: pyathena/ has no pandas/pyarrow version branches. The ImportError handling in pyathena/pandas/result_set.py, pyathena/arrow/result_set.py and pyathena/polars/result_set.py only detects installation. The pandas APIs in use (read_csv/read_parquet kwargs, TextFileReader, isetitem, infer_dtype, groupby(observed=True)) all exist in pandas 3.0.0. ParquetDataset(...).schema (pyathena/pandas/result_set.py) was only uncertain for pyarrow < 15, which the new floor excludes.
  • Tests: STRING_TYPE/STRING_NULL resolved to str/np.nan on pandas 3 (checked on 3.0.0: pd.Series(['a']).dtype.type is str, and the NULL is np.nan). The literals keep the same assertions and strictness, and the CSV branches in the same tests already used np.nan.
  • Docs: docs/null_handling.md keeps the future.infer_string opt-out, which pandas 3.0.0 and 3.0.6 still accept without a warning (string columns then become object with None).
  • Dependencies: dropping the dev numpy entries leaves numpy installed through pandas (numpy>=1.26.0, >=2.3.3 on 3.14). uv lock changed only the specifiers; the resolved versions are unchanged.

Out of scope (pre-existing, not in this diff): pyathena/arrow/util.py:97 compares type_.id with the class types.Decimal256Type instead of types.Type_DECIMAL256, so a decimal256 column maps to string. It will be filed separately if the maintainer agrees.

Limitation: the AWS floor-version runs (3.11 and 3.14) are pending.

Comment thread pyproject.toml
arrow = [
"pyarrow>=10.0.0; python_version<'3.14'",
"pyarrow>=22.0.0; python_version>='3.14'",
"pyarrow>=22.0.0",

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, callers, operations): FINDINGS (PR description only, corrected)

Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, the full PR body, commit messages, and changed docs.

Claims checked:

Existing callers: pip install -U PyAthena[pandas] or PyAthena[arrow] now upgrades pandas/pyarrow, and an environment pinned to pandas 2 or pyarrow < 22 will fail to resolve. Both are intended and listed as breaking in the release note. Installs without the extras do not check the floors, as before.

Operations: no AWS request pattern changes. The floor verification ran 485 tests per Python version (3.11 and 3.14) locally against the CI account, one run at a time, after the other Test runs had finished.

Evidence limits: the floor results are local runs at 29bfde0. The locked versions are covered only by this PR's AWS run once it is Ready. numpy at pandas' own floor and the fastparquet engine are not covered (see TEST).

Comment thread docs/null_handling.md

String columns follow the installed pandas version.
pandas 3 infers its `str` dtype and represents NULL as `NaN`; pandas 2 uses `object` columns and `None`.
String columns use the pandas `str` dtype, which represents NULL as `NaN`.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): no regressions; 1 pre-existing finding, deferred (static)

Reviewer: OpenAI Codex CLI 0.157.1 (codex exec, model gpt-6-sol, sandbox read-only, ephemeral session 01a0ea66-ae23-7773-8e08-1b51bb6a62ff), which did not author the change.
Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, in a clean detached snapshot without .env. The prompt omitted the PR number, description, commit messages, and self-review findings. The reviewer was not allowed to edit, build, test, use the network, or write to GitHub.

Reviewer's result:

Surfaces covered: The exact base-to-head diff; dependency declarations and lockfile; pandas and Arrow source and related tests; workflows; and README, introduction, pandas, Arrow, NULL-handling, and testing docs. This was static inspection only. The changed pandas tests retain the pandas 3 dtype and null assertions, and the lockfile retains NumPy through pandas.

Verdict: FINDINGS

Introduced regressions: None found.

Pre-existing issue:

  • docs/pandas.md:523: The performance examples pass unsupported read_csv options. With chunksize set, PyAthena selects the C engine even when engine="pyarrow" is requested, then forwards use_threads=True to pandas, causing the example query to fail when results are read. The nearby Python-engine low_memory and C-engine buffer_lines examples also use invalid options. This section is identical at the base commit.

Author verification: partly confirmed with pandas 3.0.6. read_csv(..., engine="c", chunksize=2, use_threads=True) and read_csv(..., buffer_lines=10) raise TypeError: read_csv() got an unexpected keyword argument, because neither is a read_csv parameter. engine="python", low_memory=True does not raise. The example is broken regardless of pandas version and is unrelated to the version floors, so it is deferred out of this PR and will be proposed as a separate docs fix.

After the review, the snapshot and PR worktree were unchanged (HEAD 29bfde0, clean status).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant