Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,8 +46,8 @@ Extra packages:
|---------------|-----------------------------------------|----------|
| SQLAlchemy | `pip install PyAthena[SQLAlchemy]` | >=2.0.0 |
| AioSQLAlchemy | `pip install PyAthena[AioSQLAlchemy]` | >=2.0.0 |
| Pandas | `pip install PyAthena[Pandas]` | >=1.3.0 |
| Arrow | `pip install PyAthena[Arrow]` | >=10.0.0 |
| Pandas | `pip install PyAthena[Pandas]` | >=3.0.0 |
| Arrow | `pip install PyAthena[Arrow]` | >=22.0.0 |
| Polars | `pip install PyAthena[Polars]` | >=1.39.0 |

## Usage
Expand Down
4 changes: 0 additions & 4 deletions docs/arrow.md
Original file line number Diff line number Diff line change
Expand Up @@ -293,10 +293,6 @@ cursor = connect(
The timeout parameters accept float values in seconds and apply to all S3 operations performed by the cursor,
including HeadObject and GetObject operations when retrieving query results.

```{note}
These timeout parameters require PyArrow >= 10.0.0, which added support for configuring S3FileSystem timeouts.
```

(async-arrow-cursor)=

## AsyncArrowCursor
Expand Down
4 changes: 2 additions & 2 deletions docs/introduction.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,8 +33,8 @@ Extra packages:
|---------------|-----------------------------------------|----------|
| SQLAlchemy | `pip install PyAthena[SQLAlchemy]` | >=2.0.0 |
| AioSQLAlchemy | `pip install PyAthena[AioSQLAlchemy]` | >=2.0.0 |
| Pandas | `pip install PyAthena[Pandas]` | >=1.3.0 |
| Arrow | `pip install PyAthena[Arrow]` | >=10.0.0 |
| Pandas | `pip install PyAthena[Pandas]` | >=3.0.0 |
| Arrow | `pip install PyAthena[Arrow]` | >=22.0.0 |
| Polars | `pip install PyAthena[Polars]` | >=1.39.0 |

(features)=
Expand Down
14 changes: 6 additions & 8 deletions docs/null_handling.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ based on actual testing with Athena:
| `Cursor` (default) | Athena API | `''` | `None` | ✅ Yes |
| `DictCursor` | Athena API | `''` | `None` | ✅ Yes |
| `PandasCursor` | CSV file | `NaN` | `NaN` | ❌ No |
| `PandasCursor` + unload | Parquet file | `''` | `NaN` (pandas 3) or `None` (pandas 2) | ✅ Yes |
| `PandasCursor` + unload | Parquet file | `''` | `NaN` | ✅ Yes |
| `ArrowCursor` | CSV file | `''` | `''` | ❌ No |
| `ArrowCursor` + unload | Parquet file | `''` | `null` | ✅ Yes |
| `PolarsCursor` | CSV file | `''` | `null` | ✅ Yes |
Expand Down Expand Up @@ -177,28 +177,26 @@ df = cursor.execute("""
print(df)
# id value description
# 0 1 empty_string <- Empty string preserved
# 1 2 NaN null_value <- NULL is NaN (None with pandas 2)
# 1 2 NaN null_value <- NULL is NaN
# 2 3 hello normal_string

print(df['value'].isna().tolist())
# [False, True, False] <- Only NULL is missing, empty string is not
```

String columns follow the installed pandas version.
pandas 3 infers its `str` dtype and represents NULL as `NaN`; pandas 2 uses `object` columns and `None`.
String columns use the pandas `str` dtype, which represents NULL as `NaN`.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): no regressions; 1 pre-existing finding, deferred (static)

Reviewer: OpenAI Codex CLI 0.157.1 (codex exec, model gpt-6-sol, sandbox read-only, ephemeral session 01a0ea66-ae23-7773-8e08-1b51bb6a62ff), which did not author the change.
Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, in a clean detached snapshot without .env. The prompt omitted the PR number, description, commit messages, and self-review findings. The reviewer was not allowed to edit, build, test, use the network, or write to GitHub.

Reviewer's result:

Surfaces covered: The exact base-to-head diff; dependency declarations and lockfile; pandas and Arrow source and related tests; workflows; and README, introduction, pandas, Arrow, NULL-handling, and testing docs. This was static inspection only. The changed pandas tests retain the pandas 3 dtype and null assertions, and the lockfile retains NumPy through pandas.

Verdict: FINDINGS

Introduced regressions: None found.

Pre-existing issue:

  • docs/pandas.md:523: The performance examples pass unsupported read_csv options. With chunksize set, PyAthena selects the C engine even when engine="pyarrow" is requested, then forwards use_threads=True to pandas, causing the example query to fail when results are read. The nearby Python-engine low_memory and C-engine buffer_lines examples also use invalid options. This section is identical at the base commit.

Author verification: partly confirmed with pandas 3.0.6. read_csv(..., engine="c", chunksize=2, use_threads=True) and read_csv(..., buffer_lines=10) raise TypeError: read_csv() got an unexpected keyword argument, because neither is a read_csv parameter. engine="python", low_memory=True does not raise. The example is broken regardless of pandas version and is unrelated to the version floors, so it is deferred out of this PR and will be proposed as a separate docs fix.

After the review, the snapshot and PR worktree were unchanged (HEAD 29bfde0, clean status).

`fetchone()`, `fetchmany()`, and `fetchall()` return the same values as the DataFrame.
Use `isna()` or `pandas.isna()` to detect NULL regardless of the pandas version.
Use `isna()` or `pandas.isna()` to detect NULL.

To get `None` for NULL strings with pandas 3, convert the columns after reading:
To get `None` for NULL strings, convert the columns after reading:

```python
df = cursor.execute("SELECT ...").as_pandas()
df = df.astype({"value": object}).where(df.notna(), None)
```

Alternatively, turn off the pandas 3 string dtype for the whole process before executing queries.
Alternatively, turn off the pandas `str` dtype for the whole process before executing queries.
String columns then use `object` with `None` for NULL, including rows returned by `fetchone()`, `fetchmany()`, and `fetchall()`.
The `future.infer_string` option exists in pandas 2.1 and later.

```python
import pandas as pd
Expand Down
14 changes: 4 additions & 10 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -56,12 +56,10 @@ aiosqlalchemy = [
"sqlalchemy[asyncio]>=2.0.0",
]
pandas = [
"pandas>=1.3.0; python_version<'3.13'",
"pandas>=2.3.0; python_version>='3.13'",
"pandas>=3.0.0",

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): FINDINGS (1, repaired)

Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 (the head of #884; this PR is stacked on it) .. head 29bfde0722f3674883960899e0eb84dc80bb7378, all 8 changed files.

Finding:

  • docs/arrow.md:296-298 (as of 2be0c42) still said the ArrowCursor connect_timeout/request_timeout options "require PyArrow >= 10.0.0". With the floor at 22.0.0, every installable version has them, so the note was obsolete. It was removed in 29bfde0. A git grep for pandas/pyarrow/numpy version mentions outside the lock finds nothing else.

Covered, no finding:

  • The library: pyathena/ has no pandas/pyarrow version branches. The ImportError handling in pyathena/pandas/result_set.py, pyathena/arrow/result_set.py and pyathena/polars/result_set.py only detects installation. The pandas APIs in use (read_csv/read_parquet kwargs, TextFileReader, isetitem, infer_dtype, groupby(observed=True)) all exist in pandas 3.0.0. ParquetDataset(...).schema (pyathena/pandas/result_set.py) was only uncertain for pyarrow < 15, which the new floor excludes.
  • Tests: STRING_TYPE/STRING_NULL resolved to str/np.nan on pandas 3 (checked on 3.0.0: pd.Series(['a']).dtype.type is str, and the NULL is np.nan). The literals keep the same assertions and strictness, and the CSV branches in the same tests already used np.nan.
  • Docs: docs/null_handling.md keeps the future.infer_string opt-out, which pandas 3.0.0 and 3.0.6 still accept without a warning (string columns then become object with None).
  • Dependencies: dropping the dev numpy entries leaves numpy installed through pandas (numpy>=1.26.0, >=2.3.3 on 3.14). uv lock changed only the specifiers; the resolved versions are unchanged.

Out of scope (pre-existing, not in this diff): pyathena/arrow/util.py:97 compares type_.id with the class types.Decimal256Type instead of types.Type_DECIMAL256, so a decimal256 column maps to string. It will be filed separately if the maintainer agrees.

Limitation: the AWS floor-version runs (3.11 and 3.14) are pending.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rebase record: patch series unchanged

After #890 and #884 merged, the branch was rebased with git rebase --onto origin/master 6aeb2ae: old 6aeb2ae..29bfde0722f3674883960899e0eb84dc80bb7378, new 6effaa3..842f0f05b1ac8dc456f94b35d390e89a85defc70. The published head was replaced with --force-with-lease against 29bfde0, and the base is now master.

  • git range-diff 6aeb2ae..29bfde0 origin/master..842f0f0: both commits = (identical).
  • Upstream changes since 6aeb2ae: Leave unset storage values out of the Glue table parameters #890 (pyathena/glue.py, tests/pyathena/test_glue.py) and fix: render Hive STRUCT syntax in table column DDL #870 (pyathena/sqlalchemy/compiler.py, docs/sqlalchemy.md, SQLAlchemy tests). Neither touches the pandas/Arrow/Polars code paths, the dependency declarations, or the files this PR changes, so the review rounds and the floor-version runs at 29bfde0 still apply.
  • At 842f0f0: uv lock --check resolves without changes, and just lint and just docs lint pass.

The #890 fix is now in the base, so the Ready run is expected to clear the Glue metadata failures seen in run 36499867658.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ready run at 842f0f0: all passed

The Ready run 36589928431 passed test, test-sqla, and test-sqla-async on Python 3.14, including the Glue metadata tests that failed in 36499867658. Offline checks and the docs build also pass, and GitHub reports the PR as mergeable (clean).

After this run, master gained #883 (sync/aio cursor base dedupe) and #895 (SQLAlchemy type_compiler_cls). Neither overlaps this PR's files. In the pandas/Arrow/Polars cursors, #883 only renames the base mixin (WithFetch → WithResultSet), which does not interact with the dependency floors or with the changed test constants. So the branch was not rebased again.

]
arrow = [
"pyarrow>=10.0.0; python_version<'3.14'",
"pyarrow>=22.0.0; python_version>='3.14'",
"pyarrow>=22.0.0",

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, callers, operations): FINDINGS (PR description only, corrected)

Scope: base 6aeb2ae9967f378478737cf6bcd7113c3f4342b6 .. head 29bfde0722f3674883960899e0eb84dc80bb7378, the full PR body, commit messages, and changed docs.

Claims checked:

Existing callers: pip install -U PyAthena[pandas] or PyAthena[arrow] now upgrades pandas/pyarrow, and an environment pinned to pandas 2 or pyarrow < 22 will fail to resolve. Both are intended and listed as breaking in the release note. Installs without the extras do not check the floors, as before.

Operations: no AWS request pattern changes. The floor verification ran 485 tests per Python version (3.11 and 3.14) locally against the CI account, one run at a time, after the other Test runs had finished.

Evidence limits: the floor results are local runs at 29bfde0. The locked versions are covered only by this PR's AWS run once it is Ready. numpy at pandas' own floor and the fastparquet engine are not covered (see TEST).

]
polars = [
"polars>=1.39.0",
Expand All @@ -70,12 +68,8 @@ polars = [
[dependency-groups]
dev = [
"sqlalchemy[asyncio]>=2.0.0",
"pandas>=1.3.0; python_version<'3.13'",
"pandas>=2.3.0; python_version>='3.13'",
"numpy>=1.26.0; python_version<'3.13'",
"numpy>=2.3.0; python_version>='3.14'",
"pyarrow>=10.0.0; python_version<'3.14'",
"pyarrow>=22.0.0; python_version>='3.14'",
"pandas>=3.0.0",
"pyarrow>=22.0.0",
"polars>=1.39.0",
"Jinja2>=3.1.0",
"mypy>=0.900",
Expand Down
9 changes: 2 additions & 7 deletions tests/pyathena/pandas/test_async_cursor.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,6 @@
from tests import ENV
from tests.pyathena.conftest import connect

# pandas 3 infers its "str" dtype for strings, which represents NULL as NaN; pandas 2 uses
# object columns with None.
STRING_TYPE = pd.Series(["a"]).dtype.type
STRING_NULL = pd.Series(["a", None]).iloc[1]


class TestAsyncPandasCursor:
def test_binary_null_vs_empty(self, async_pandas_cursor):
Expand Down Expand Up @@ -592,7 +587,7 @@ def test_empty_and_null_string(self, async_pandas_cursor, parquet_engine):
# NULL and empty characters are correctly converted when the UNLOAD option is enabled.
np.testing.assert_equal(
result_set.fetchall(),
[("", "a"), ("N/A", "a"), ("NULL", "a"), (STRING_NULL, "a")],
[("", "a"), ("N/A", "a"), ("NULL", "a"), (np.nan, "a")],
)
else:
np.testing.assert_equal(
Expand All @@ -605,7 +600,7 @@ def test_empty_and_null_string(self, async_pandas_cursor, parquet_engine):
# NULL and empty characters are correctly converted when the UNLOAD option is enabled.
np.testing.assert_equal(
result_set.fetchall(),
[("", "a"), ("N/A", "a"), ("NULL", "a"), (STRING_NULL, "a")],
[("", "a"), ("N/A", "a"), ("NULL", "a"), (np.nan, "a")],
)
else:
assert result_set.fetchall() == [
Expand Down
23 changes: 9 additions & 14 deletions tests/pyathena/pandas/test_cursor.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,11 +20,6 @@
from tests import ENV
from tests.pyathena.conftest import connect

# pandas 3 infers its "str" dtype for strings, which represents NULL as NaN; pandas 2 uses
# object columns with None.
STRING_TYPE = pd.Series(["a"]).dtype.type
STRING_NULL = pd.Series(["a", None]).iloc[1]


class TestPandasCursor:
@pytest.mark.parametrize(
Expand Down Expand Up @@ -646,17 +641,17 @@ def test_complex_as_pandas(self, pandas_cursor, chunksize):
np.int64,
np.float64,
np.float64,
STRING_TYPE,
STRING_TYPE,
str,
str,
np.datetime64,
np.object_,
np.datetime64,
np.object_,
STRING_TYPE,
str,
np.object_,
STRING_TYPE,
str,
np.object_,
STRING_TYPE,
str,
np.object_,
)
rows = [
Expand Down Expand Up @@ -768,8 +763,8 @@ def test_complex_unload_as_pandas_pyarrow(self, pandas_cursor, parquet_engine):
np.int64,
np.float32,
np.float64,
STRING_TYPE,
STRING_TYPE,
str,
str,
np.datetime64,
np.object_,
np.object_,
Expand Down Expand Up @@ -1200,7 +1195,7 @@ def test_null_vs_empty_string(self, pandas_cursor, parquet_engine):
# NULL and empty characters are correctly converted when the UNLOAD option is enabled.
np.testing.assert_equal(
pandas_cursor.fetchall(),
[("", "a"), ("N/A", "a"), ("NULL", "a"), (STRING_NULL, "a")],
[("", "a"), ("N/A", "a"), ("NULL", "a"), (np.nan, "a")],
)
else:
np.testing.assert_equal(
Expand All @@ -1212,7 +1207,7 @@ def test_null_vs_empty_string(self, pandas_cursor, parquet_engine):
# NULL and empty characters are correctly converted when the UNLOAD option is enabled.
np.testing.assert_equal(
pandas_cursor.fetchall(),
[("", "a"), ("N/A", "a"), ("NULL", "a"), (STRING_NULL, "a")],
[("", "a"), ("N/A", "a"), ("NULL", "a"), (np.nan, "a")],
)
else:
assert pandas_cursor.fetchall() == [
Expand Down
20 changes: 6 additions & 14 deletions uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading