Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
0fac8b1
Backport #769: Enable Iceberg UPDATE target subquery compliance
laughingman7743 Sep 29, 2026
7592993
Backport #772: Restore SQLAlchemy CTE compliance coverage
laughingman7743 Sep 29, 2026
311ce59
Backport #770: Restore SQLAlchemy binary compliance and preserve CSV …
laughingman7743 Sep 29, 2026
048498b
Backport #789: Skip AWS integration jobs for external fork pull requests
laughingman7743 Sep 29, 2026
27331ec
Backport #774: Support native SQLAlchemy ARRAY types and typed round …
laughingman7743 Sep 29, 2026
5e86f01
Backport #775: Support SQLAlchemy ARRAY expressions in SELECT and WHERE
laughingman7743 Sep 29, 2026
e834e92
Backport #776: Support SQLAlchemy ARRAY partial UPDATE and slice resi…
laughingman7743 Sep 29, 2026
b43504f
Backport #811: Rerun tests once on Athena service-side query failures
laughingman7743 Sep 29, 2026
2290413
Backport #815: Give each test session its own S3 Tables namespace
laughingman7743 Sep 29, 2026
29a4d0a
Backport #814: Run the three test suites in parallel
laughingman7743 Sep 29, 2026
c521180
Backport #825: Re-enable SQLAlchemy parameter and SQL rendering compl…
laughingman7743 Sep 29, 2026
c4dfbbd
Backport #826: SQLAlchemy compliance: audit numeric range and float p…
laughingman7743 Sep 29, 2026
70a763d
Backport #837: Run AWS test suites only on ready pull requests and re…
laughingman7743 Sep 29, 2026
91c922b
Backport #863: Run pull-request AWS tests on the newest Python versio…
laughingman7743 Sep 29, 2026
8340b4e
Backport #862: Run the full test matrix in the Release workflow befor…
laughingman7743 Sep 29, 2026
caaae8f
Backport #866: Set up the PyAthena test fixtures only in processes th…
laughingman7743 Sep 29, 2026
f0380bf
Backport #867: Insert SQLAlchemy compliance fixture rows once for rea…
laughingman7743 Sep 29, 2026
08645a1
Backport #873: Give the executemany and partition tests their own tables
laughingman7743 Sep 29, 2026
3916345
Backport #870: fix: render Hive STRUCT syntax in table column DDL
laughingman7743 Sep 29, 2026
e3ec1a9
Keep the ARRAY support working on SQLAlchemy 1.x (3.x only)
laughingman7743 Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 7 additions & 3 deletions .github/workflows/docs-trigger.yaml
Original file line number Diff line number Diff line change
@@ -1,14 +1,18 @@
name: Trigger Docs on Tag
name: Trigger Docs on Release

# Rebuilds the documentation after the Release workflow succeeds, so a tag it
# refuses to publish does not trigger a documentation build.
on:
push:
tags: ['v*']
workflow_run:
workflows: [Release]
types: [completed]

permissions:
actions: write

jobs:
trigger-docs:
if: github.event.workflow_run.conclusion == 'success'
runs-on: ubuntu-latest
steps:
- uses: actions/github-script@ed597411d8f924073f98dfc5c65a23a2325f34cd # v8.0.0
Expand Down
17 changes: 17 additions & 0 deletions .github/workflows/docs.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,23 @@ jobs:
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0 # Fetch all history for sphinx-multiversion
# sphinx-multiversion builds every version tag in the checkout. A tag
# gets its GitHub release only after the Release workflow's tests and
# PyPI upload succeed, so drop tags without one: a release still in
# progress or refused by its tests is not documented.
- name: Drop unreleased version tags
env:
GH_TOKEN: ${{ github.token }}
run: |
released=$(gh api --paginate "repos/$GITHUB_REPOSITORY/releases?per_page=100" \
--jq '.[] | select(.draft | not) | .tag_name')
if [[ -z "$released" ]]; then
echo "::error::No published GitHub releases found"
exit 1
fi
git tag --list 'v*' | while read -r tag; do
grep -qxF "$tag" <<< "$released" || git tag --delete "$tag"
done
- name: Setup Pages
uses: actions/configure-pages@45bfe0192ca1faeb007ade9deae92b16b8254a0d # v6.0.0
- uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7.6.0
Expand Down
17 changes: 14 additions & 3 deletions .github/workflows/release.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,24 @@ on:
tags:
- 'v*'

permissions:
id-token: write
contents: write
permissions: {}

jobs:
# Runs every suite on every supported Python version for the tagged commit;
# nothing is built or published unless all of them pass.
test:
uses: ./.github/workflows/test.yaml
permissions:
contents: read
id-token: write
pull-requests: read

release:
needs: test

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, callers, and operations), by the authoring model; not an independent review.

Scope: the same base and head as round one (4ef4d3e..e3ec1a9). This round covers the PR body, the commit messages, and the docs added to 3.x.

Result: FINDINGS. Two gaps in the PR body, both corrected; no code change.

  1. The release notes omitted Support native SQLAlchemy ARRAY types and typed round trips #774's compatibility limits for SQLAlchemy queries that select ARRAY values. Rewritten queries need explicit SELECT columns and labeled literal expressions. Textual ORDER BY is limited to selected names, lists of them, and ordinals. Decimal ARRAY binds need an explicit precision. These are behavior changes for a minor release, so they are now listed, in Support native SQLAlchemy ARRAY types and typed round trips #774's own wording.
  2. The list of conflict sources omitted Manage the benchmark project as a uv workspace member #813, the source of the benchmarks/uv.lock conflict. It is now listed.

Claims checked:

  • The per-PR conflict table matches the commit messages.
  • The release-note items match the code:
    • array in ischema_names and types.ARRAY in colspecs map to AthenaArray.
    • STRUCT and INT render in column DDL.
    • CAST to Double renders DOUBLE.
  • The SQLAlchemy 1.x statement matches the recorded 1.4.54/2.0.46 check, including its limit: compile level only.
  • Docs: the floating-point table matches 3.x's compiler. The CI-policy text in docs/testing.md describes the repository's Test and Release workflows, and it holds for 3.x after this PR. The scheduled row applies to the default branch, because GitHub runs schedules only from there.

Existing callers:

  • Python 3.10 remains supported: master supported 3.10 when these PRs merged, and the release gate tests 3.10 through 3.14.
  • SQLAlchemy 1.x keeps compiling the STRUCT, MAP, and ARRAY casts it compiled before.

Operations:

  • A ready 3.x PR now runs the suites on the newest Python version, and only the related suites on later pushes.
  • A tag runs every suite on five Python versions before publishing. With Rerun tests once on Athena service-side query failures #811, an Athena service-side failure is rerun once. A remaining failure blocks publishing until the failed jobs are rerun, or the tag is recreated.
  • docs-trigger now fires from the default branch's workflow after a successful Release. The tag's own copy no longer starts a docs build.

runs-on: ubuntu-latest
permissions:
id-token: write
contents: write

env:
PYTHON_VERSION: '3.12'
Expand Down
27 changes: 22 additions & 5 deletions .github/workflows/test-suite.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,28 +6,45 @@ on:
test-type:
required: true
type: string
python-versions:
description: JSON array of the Python versions to test
required: true
type: string
skip-spark:
description: Skip the Spark tests of the PyAthena suite
type: boolean
default: false
skip-sqla:
description: Skip the SQLAlchemy tests of the PyAthena suite
type: boolean
default: false

jobs:
run:
# External fork contributions must validate AWS behavior in their own account.
if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest

env:
TEST_TYPE: ${{ inputs.test-type }}
PYTEST_ADDOPTS: >-
${{ inputs.skip-spark && '--ignore=tests/pyathena/spark --ignore=tests/pyathena/aio/spark' || '' }}
${{ inputs.skip-sqla && '--ignore=tests/pyathena/sqlalchemy --ignore=tests/pyathena/aio/sqlalchemy' || '' }}
AWS_DEFAULT_REGION: us-west-2
AWS_ATHENA_S3_STAGING_DIR: s3://laughingman7743-pyathena/github/
AWS_ATHENA_WORKGROUP: pyathena
AWS_ATHENA_SPARK_WORKGROUP: pyathena-spark
AWS_ATHENA_MANAGED_WORKGROUP: pyathena-managed
# Registered S3 Tables catalog (s3tablescatalog/<table-bucket>) and namespace
# for the SQLAlchemy S3 Tables tests; the table bucket, namespace, and the
# AWS analytics-services integration are provisioned out of band.
# Registered S3 Tables catalog (s3tablescatalog/<table-bucket>) for the S3
# Tables tests; each test session creates and deletes its own namespace.
# The table bucket and the AWS analytics-services integration are
# provisioned out of band.
AWS_ATHENA_S3_TABLES_CATALOG: s3tablescatalog/laughingman7743-pyathena-s3-tables
AWS_ATHENA_S3_TABLES_NAMESPACE: pyathena

strategy:
fail-fast: false
matrix:
python-version: ['3.10', '3.11', '3.12', '3.13', '3.14']
python-version: ${{ fromJSON(inputs.python-versions) }}

steps:
- name: Checkout
Expand Down
126 changes: 121 additions & 5 deletions .github/workflows/test.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,30 @@ name: Test

on:
pull_request:
# ready_for_review starts the AWS jobs for a pull request leaving Draft;
# converted_to_draft starts a run without them, which cancels an
# in-progress run through the concurrency group. paths-ignore applies to
# both, so a pull request that now changes only docs starts neither.
types: [opened, synchronize, reopened, ready_for_review, converted_to_draft]
paths-ignore:
- 'docs/**'
- '**.md'
# The scheduled run executes every suite, including the ones that pull
# requests only run when related files change, on the newest Python version.
schedule:
- cron: '0 0 * * 0'
# Allows refreshing the README status badge on demand: the badge reflects
# the latest run on the default branch, which is otherwise only the weekly
# scheduled run and stays red for up to a week after a transient failure.
# Runs every suite on the selected branch: on demand for a pull request, and
# to refresh the README status badge after a transient failure on the
# default branch.
workflow_dispatch:
inputs:
python-versions:
description: Comma-separated Python versions, such as 3.12 or 3.11,3.14; empty for every supported version
type: string
default: ''
# The Release workflow runs every suite on every supported Python version
# before publishing.
workflow_call:

permissions:
id-token: write
Expand All @@ -23,19 +38,120 @@ concurrency:
cancel-in-progress: true

jobs:
# Offline checks run for every event, including Draft and fork pull requests.
lint:
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7.6.0
with:
python-version: '3.12'
enable-cache: true
- uses: taiki-e/install-action@7a79fe8c3a13344501c80d99cae481c1c9085912 # v2.81.10
with:
tool: just
- run: just lint

# Selects the AWS suites and Python versions. Draft and external-fork pull
# requests run none. A ready pull request always runs the PyAthena suite; it
# runs the SQLAlchemy tests (the compliance suites and the PyAthena suite's
# SQLAlchemy tests) and the Spark tests only when their code, tests,
# dependencies, or this workflow change. Pull requests and the schedule test
# the newest Python version; a dispatch tests the requested versions or
# every version, and the Release workflow every version.
changes:
if: >-
github.event_name != 'pull_request' ||
(!github.event.pull_request.draft &&
github.event.pull_request.head.repo.full_name == github.repository)
runs-on: ubuntu-latest
permissions:
pull-requests: read
outputs:
sqla: ${{ steps.filter.outputs.sqla }}
spark: ${{ steps.filter.outputs.spark }}
python-versions: ${{ steps.filter.outputs.python-versions }}
steps:
- id: filter
env:
GH_TOKEN: ${{ github.token }}
EVENT_NAME: ${{ github.event_name }}
REPO: ${{ github.repository }}
PR_NUMBER: ${{ github.event.pull_request.number }}
REQUESTED_VERSIONS: ${{ inputs.python-versions }}
# Every supported version, oldest first; keep in sync with the
# pyproject.toml classifiers.
PYTHON_VERSIONS: '["3.10", "3.11", "3.12", "3.13", "3.14"]'
run: |
case "$EVENT_NAME" in
pull_request | schedule)
versions=$(jq -c '[last]' <<< "$PYTHON_VERSIONS")
;;
workflow_dispatch)
versions=$(jq -c --arg requested "$REQUESTED_VERSIONS" '
($requested | split(",") | map(gsub("\\s"; "")) | map(select(. != "")) | unique) as $selected
| if $selected == [] then .
elif ($selected - .) == [] then $selected
else error("unsupported Python versions: \($selected - . | join(", "))")
end' <<< "$PYTHON_VERSIONS")
;;
*)
# The Release workflow (a workflow_call from a tag push).
versions=$(jq -c '.' <<< "$PYTHON_VERSIONS")
;;
esac
echo "python-versions=$versions" >> "$GITHUB_OUTPUT"
if [[ "$EVENT_NAME" != "pull_request" ]]; then
{
echo "sqla=true"
echo "spark=true"
} >> "$GITHUB_OUTPUT"
exit 0
fi
files=$(gh api "repos/$REPO/pulls/$PR_NUMBER/files" --paginate --jq '.[].filename')
printf 'Changed files:\n%s\n' "$files"
shared='^(\.github/workflows/test(-suite)?\.yaml|justfile|pyproject\.toml|uv\.lock)$'
sqla="$shared|^pyathena/(aio/)?sqlalchemy/|^tests/sqlalchemy/|^tests/pyathena/(aio/)?sqlalchemy/|^setup\.cfg$"

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review, relayed. The reviewer was Codex CLI 0.157.1 (codex exec --sandbox read-only), a different model from the author.

Scope: 4ef4d3e..e3ec1a9 in a detached snapshot. The prompt gave the head, the source merge SHAs, and the scope; it did not include the PR number, PR body, commit messages, or earlier findings. Static review only: no tests, builds, GitHub access, or network.

Covered: all 19 cherry-picks compared with their source merges, including the conflict resolutions; the 3.x-only SQLAlchemy 1.x change; affected cursors and dialect code; tests; docs; Actions workflows.

Result: FINDINGS.

  • "The backported code matches the source changes apart from the documented 3.x adjustments. I found no non-backported master changes leaking into the diff."
  • P2: the changes path filter here omits pyathena/util.py, which the SQLAlchemy and Spark packages use. A ready PR that changes only a shared module would skip the SQLAlchemy and Spark tests.
  • It noted, without counting it as a finding, that namespaces left by interrupted runs are recovered only by the master-only sweep.

Verification and disposition:

spark="$shared|^pyathena/(aio/)?spark/|^tests/pyathena/(aio/)?spark/"
if grep -qE "$sqla" <<< "$files"; then
echo "sqla=true" >> "$GITHUB_OUTPUT"
else
echo "sqla=false" >> "$GITHUB_OUTPUT"
fi
if grep -qE "$spark" <<< "$files"; then
echo "spark=true" >> "$GITHUB_OUTPUT"
else
echo "spark=false" >> "$GITHUB_OUTPUT"
fi

# The three suites create their own schemas and tables, so they run in
# parallel; each is still a separate job for "Re-run failed jobs".
test:
needs: changes
uses: ./.github/workflows/test-suite.yaml
with:
test-type: pyathena
python-versions: ${{ needs.changes.outputs.python-versions }}
skip-spark: ${{ needs.changes.outputs.spark != 'true' }}
skip-sqla: ${{ needs.changes.outputs.sqla != 'true' }}

test-sqla:
needs: [test]
needs: changes
if: needs.changes.outputs.sqla == 'true'
uses: ./.github/workflows/test-suite.yaml
with:
test-type: sqla
python-versions: ${{ needs.changes.outputs.python-versions }}

test-sqla-async:
needs: [test-sqla]
needs: changes
if: needs.changes.outputs.sqla == 'true'
uses: ./.github/workflows/test-suite.yaml
with:
test-type: sqla_async
python-versions: ${{ needs.changes.outputs.python-versions }}
11 changes: 11 additions & 0 deletions docs/null_handling.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,17 @@ which correctly interprets unquoted empty values as NULL, while `S3FSCursor` use
`AthenaCSVReader` that respects CSV quoting rules.
```

## Binary Values

The string comparison above does not apply to `VARBINARY` columns.
With the default CSV settings and converters, pandas and Arrow cursors distinguish SQL NULL from empty binary
values when reading CSV results: `fetchone()`, `fetchmany()`, and `fetchall()` return `None`
for NULL and `b''` for an empty binary value.
This also applies to their asynchronous variants and pandas chunked reads.
Pandas DataFrames preserve the same values.
Arrow Tables retain the CSV hexadecimal strings, with NULL represented as an Arrow null;
fetch methods convert the hexadecimal strings to Python bytes.

## Default Cursor (API-based)

The default `Cursor` and `DictCursor` fetch results directly from the Athena API,
Expand Down
Loading
Loading