Require pandas 3.0 and pyarrow 22.0 - #885
laughingman7743 wants to merge 2 commits into
Conversation
The pandas and arrow extras declared floors that CI never installed (pandas>=1.3.0, pyarrow>=10.0.0), so they promised compatibility that was not verified. With Python 3.10 dropped, the extras can require the versions the project tests: - pandas>=3.0.0, the first release of the only major version CI tests. - pyarrow>=22.0.0, the first release with wheels for every supported Python version, including 3.14. The dev group follows the extras and no longer lists numpy, whose markers only restated pandas' own numpy requirement. The pandas cursor tests and the NULL handling guide drop their pandas 2 branches. Closes #852 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every supported pyarrow version accepts the S3FileSystem timeouts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| pandas = [ | ||
| "pandas>=1.3.0; python_version<'3.13'", | ||
| "pandas>=2.3.0; python_version>='3.13'", | ||
| "pandas>=3.0.0", |
There was a problem hiding this comment.
Self-review round one (implementation behavior): FINDINGS (1, repaired)
Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 (the head of #884; this PR is stacked on it) .. head 29bfde0722f3674883960899e0eb84dc80bb7378, all 8 changed files.
Finding:
docs/arrow.md:296-298(as of2be0c42) still said the ArrowCursorconnect_timeout/request_timeoutoptions "require PyArrow >= 10.0.0". With the floor at 22.0.0, every installable version has them, so the note was obsolete. It was removed in29bfde0. Agit grepfor pandas/pyarrow/numpy version mentions outside the lock finds nothing else.
Covered, no finding:
- The library:
pyathena/has no pandas/pyarrow version branches. TheImportErrorhandling inpyathena/pandas/result_set.py,pyathena/arrow/result_set.pyandpyathena/polars/result_set.pyonly detects installation. The pandas APIs in use (read_csv/read_parquetkwargs,TextFileReader,isetitem,infer_dtype,groupby(observed=True)) all exist in pandas 3.0.0.ParquetDataset(...).schema(pyathena/pandas/result_set.py) was only uncertain for pyarrow < 15, which the new floor excludes. - Tests:
STRING_TYPE/STRING_NULLresolved tostr/np.nanon pandas 3 (checked on 3.0.0:pd.Series(['a']).dtype.type is str, and the NULL isnp.nan). The literals keep the same assertions and strictness, and the CSV branches in the same tests already usednp.nan. - Docs:
docs/null_handling.mdkeeps thefuture.infer_stringopt-out, which pandas 3.0.0 and 3.0.6 still accept without a warning (string columns then becomeobjectwithNone). - Dependencies: dropping the dev
numpyentries leaves numpy installed through pandas (numpy>=1.26.0,>=2.3.3on 3.14).uv lockchanged only the specifiers; the resolved versions are unchanged.
Out of scope (pre-existing, not in this diff): pyathena/arrow/util.py:97 compares type_.id with the class types.Decimal256Type instead of types.Type_DECIMAL256, so a decimal256 column maps to string. It will be filed separately if the maintainer agrees.
Limitation: the AWS floor-version runs (3.11 and 3.14) are pending.
| arrow = [ | ||
| "pyarrow>=10.0.0; python_version<'3.14'", | ||
| "pyarrow>=22.0.0; python_version>='3.14'", | ||
| "pyarrow>=22.0.0", |
There was a problem hiding this comment.
Self-review round two (claims, callers, operations): FINDINGS (PR description only, corrected)
Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, the full PR body, commit messages, and changed docs.
Claims checked:
- "Polars before 1.39 silently truncated results" (the Require polars>=1.39.0 for chunked PolarsCursor reads #828 example, repeated from the Raise the pandas and pyarrow minimum versions to tested versions #852 premise): the Require polars>=1.39.0 for chunked PolarsCursor reads #828 measurements show that only Polars 1.34–1.36.x ended a failed chunked read without an error (PolarsCursor with chunksize returns 0 rows instead of raising when the CSV read fails #820). Before 1.34 it failed loudly, and 1.37–1.38 had the file-cache issue (PolarsCursor chunked CSV reads download the whole result into Polars' local file cache #821), not truncation. The PR body now says "with Polars 1.34–1.36, a chunked read that failed partway ended without an error (PolarsCursor with chunksize returns 0 rows instead of raising when the CSV read fails #820)".
- "pyarrow 22.0.0 (2025-10-24) is the first release with wheels for every supported Python version; 21.0.0 has no cp314 wheels": PyPI file lists show 21.0.0 wheels for cp39–cp313 and 22.0.0 for cp310–cp314, with 22.0.0 uploaded 2025-10-24.
- "pandas 3.0 requires
numpy>=1.26.0(>=2.3.3on 3.14)" and "pyarrow 13.0.0 for its pyarrow features": the pandas 3.0.0requires_diston PyPI. The removed dev markers (<'3.13',>='3.14') did leave 3.13 unconstrained. - The
future.infer_stringopt-out: pandas 3.0.0 and 3.0.6 acceptFalsewithout a warning and returnobject/None. - "String columns use the pandas
strdtype, NULL isNaN" (docs/null_handling.md): Parquet reads go through the pyarrow engine only (_get_parquet_engine), andtest_null_vs_empty_stringassertsnp.nanfor theautoandpyarrowengines. The floor runs passed it on 3.11 and 3.14. - "
uv.lockonly changes specifiers": the diff has only specifier and marker lines; noversion =line changes.
Existing callers: pip install -U PyAthena[pandas] or PyAthena[arrow] now upgrades pandas/pyarrow, and an environment pinned to pandas 2 or pyarrow < 22 will fail to resolve. Both are intended and listed as breaking in the release note. Installs without the extras do not check the floors, as before.
Operations: no AWS request pattern changes. The floor verification ran 485 tests per Python version (3.11 and 3.14) locally against the CI account, one run at a time, after the other Test runs had finished.
Evidence limits: the floor results are local runs at 29bfde0. The locked versions are covered only by this PR's AWS run once it is Ready. numpy at pandas' own floor and the fastparquet engine are not covered (see TEST).
|
|
||
| String columns follow the installed pandas version. | ||
| pandas 3 infers its `str` dtype and represents NULL as `NaN`; pandas 2 uses `object` columns and `None`. | ||
| String columns use the pandas `str` dtype, which represents NULL as `NaN`. |
There was a problem hiding this comment.
Independent review (relayed): no regressions; 1 pre-existing finding, deferred (static)
Reviewer: OpenAI Codex CLI 0.157.1 (codex exec, model gpt-6-sol, sandbox read-only, ephemeral session 01a0ea66-ae23-7773-8e08-1b51bb6a62ff), which did not author the change.
Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, in a clean detached snapshot without .env. The prompt omitted the PR number, description, commit messages, and self-review findings. The reviewer was not allowed to edit, build, test, use the network, or write to GitHub.
Reviewer's result:
Surfaces covered: The exact base-to-head diff; dependency declarations and lockfile; pandas and Arrow source and related tests; workflows; and README, introduction, pandas, Arrow, NULL-handling, and testing docs. This was static inspection only. The changed pandas tests retain the pandas 3 dtype and null assertions, and the lockfile retains NumPy through pandas.
Verdict: FINDINGS
Introduced regressions: None found.
Pre-existing issue:
- docs/pandas.md:523: The performance examples pass unsupported
read_csvoptions. Withchunksizeset, PyAthena selects the C engine even whenengine="pyarrow"is requested, then forwardsuse_threads=Trueto pandas, causing the example query to fail when results are read. The nearby Python-enginelow_memoryand C-enginebuffer_linesexamples also use invalid options. This section is identical at the base commit.
Author verification: partly confirmed with pandas 3.0.6. read_csv(..., engine="c", chunksize=2, use_threads=True) and read_csv(..., buffer_lines=10) raise TypeError: read_csv() got an unexpected keyword argument, because neither is a read_csv parameter. engine="python", low_memory=True does not raise. The example is broken regardless of pandas version and is unrelated to the version floors, so it is deferred out of this PR and will be proposed as a separate docs fix.
After the review, the snapshot and PR worktree were unchanged (HEAD 29bfde0, clean status).
WHAT
Raise the pandas and Arrow extras to versions the project tests, for 4.0.0. Stacked on #884 (drop Python 3.10); the base will be retargeted to
masteronce #884 merges.pandasextra:pandas>=3.0.0(was>=1.3.0below Python 3.13,>=2.3.0from 3.13).arrowextra:pyarrow>=22.0.0for every Python version (was>=10.0.0below Python 3.14).numpyentries. pandas 3.0 requiresnumpy>=1.26.0(>=2.3.3on Python 3.14) itself, and the removed markers only restated that. They also left Python 3.13 unconstrained. The tests still import numpy, which pandas installs.README.mdanddocs/introduction.mdlist the new minimums.docs/arrow.mddrops a note saying the ArrowCursor S3 timeouts need PyArrow 10.0.0 or later, since every supported version has them.docs/null_handling.mdand the pandas cursor tests (tests/pyathena/pandas/test_cursor.py,test_async_cursor.py) drop their pandas 2 branches. String columns are the pandasstrdtype, and NULL isNaN. Thefuture.infer_stringopt-out stays documented because pandas 3.0.0 and 3.0.6 still accept it.uv.lock: only the declared specifiers change; the resolved versions do not.The library itself has no pandas or pyarrow version branches to remove. Its
ImportErrorhandling only detects whether an optional package is installed.Release note (breaking): the
pandasextra requires pandas 3.0.0 or later, and thearrowextra requires pyarrow 22.0.0 or later. Users who need pandas 2 or an older pyarrow can stay on the 3.x maintenance branch.WHY
Closes #852. The declared floors were never installed by CI, so they promised compatibility that was not verified. In #828, an untested floor hid a real defect: with Polars 1.34–1.36, a chunked read that failed partway ended without an error (#820).
TEST
Tested commit: 29bfde0
just format,just lint: pass.just docs lint: pass.just docs build: succeeds, with no warnings from the changed pages.UV_PROJECT_ENVIRONMENT=<venv> uv run --no-sync --env-file .env python -m pytest -n 8 --reruns 1 --only-rerun … tests/pyathena/{pandas,arrow,polars} tests/pyathena/aio/{pandas,arrow,polars} tests/pyathena/sqlalchemy/test_base.py::TestSQLAlchemyAthena::{test_to_sql_parquet,test_to_sql_json,test_to_sql_column_options}.ParserWarnings fromtest_binary_custom_dialect, which passes a CSV dialect whosequotingoverrides PyAthena's.Not covered: numpy at pandas' own floor (1.26.0), and the fastparquet engine, which is not a dependency of any extra.
🤖 Generated with Claude Code