Skip to content

Define the shared test tables in Python and generate their data - #874

Merged
laughingman7743 merged 2 commits into
masterfrom
test/848-python-fixture-tables
Sep 28, 2026
Merged

laughingman7743 merged 2 commits into
masterfrom
test/848-python-fixture-tables

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

WHAT

  • New tests/pyathena/tables.py defines the shared test tables and views once, in Python:
    • Column(name, athena_type, arrow_type, comment)
    • Table(name, columns, rows, storage, partitions, comment, tblproperties), where rows are Python values in column order and storage is "parquet" or "text"
    • View(name, query)
    • TABLES and VIEWS hold the same tables and views that create_table.sql.jinja2 created: one_row, many_rows, one_row_complex, partition_table, integer_na_values, boolean_na_values, parquet_with_compression, view_one_row, v_one_row.
    • Table.create_statement generates the DDL, and Table.data_file generates the data file: Parquet written with pyarrow using the column's Arrow type, or tab-separated text.
    • The Spark tests' spark_group_by.csv is generated from SPARK_GROUP_BY.
  • tests/pyathena/conftest.py: the session setup uploads the generated files (_upload_data, _delete_data) and creates the tables and views from the definitions (_create_tables). The data files go to <schema>/<table>/data.parquet or data.tsv. The Spark CSV and the filesystem test file keep their keys.
  • Removed: tests/resources/rows/ (6 files) and tests/resources/queries/create_table.sql.jinja2, together with its entry in scripts/config/license_headers.toml.

Storage:

Table Before After Why
one_row text text The reflection tests assert its LazySimpleSerDe, text input format, and field.delim/line.delim SerDe properties
many_rows text text Tests read it without ORDER BY and expect file order. With a Parquet file, SELECT * FROM many_rows LIMIT 100 returned rows out of order (0..6, 15..30, 7..), which failed test_pandas_cursor_chunked_vs_regular_same_data and test_pandas_cursor_iter_chunks_consistency
one_row_complex, integer_na_values, boolean_na_values text (gzipped Hive text for one_row_complex) Parquet Nested, binary, decimal, and NULL values are written as values, with no delimiter encoding
partition_table, parquet_with_compression no data no data partition_table changes from text to Parquet; the tests only read its metadata

WHY

Part of #848 (step 1), for #834. Stacked on #873.

one_row_complex.gz was a single gzipped Hive text row with \002/\003 separators, which could not practically be edited by hand. Adding a column meant editing that file and the DDL template separately. Each table is now one entry, and the DDL and data are generated from it. Deriving the test expectations from these definitions is the next step of #848.

The session runs the same statements as before: 7 CREATE EXTERNAL TABLE and 2 CREATE VIEW per test process. It uploads the same number of objects, 7 per process: 5 data files, the Spark CSV, and the filesystem test file.

TEST

Tested commit: 12cfd65. The only later commit types Table.storage as Literal["parquet", "text"], with no change in behavior; just lint passed on it.

  • just lint passed.
  • Generated text files compared offline with the removed files: one_row and many_rows data and spark_group_by.csv are byte-identical.
  • Parity against AWS: a script, kept outside the repository, created the old tables from the removed template and data files in one schema and the generated tables in another. It compared SELECT * (all rows, in order) and DESCRIBE for all 7 tables and 2 views, and found no differences. This ran while many_rows was still Parquet, and even so its 10,000 rows came back in file order that time. many_rows is now text again, byte-identical to the old file.
  • Against AWS, from this worktree:
    • pytest -n 4 tests/pyathena -k "complex or na_values or reflect or table_options or get_columns or table_comment or view or glue or as_pandas or partition or spark_dataframe or spark_sql or table_names or has_table or chunk": 282 passed, 2 skipped. The 2 failures were the many_rows ordering failures above, which led to keeping it as text.
    • After that change, pytest -n 4 tests/pyathena -k "many or fetchmany or fetchall or iter or chunk or arraysize": 371 passed, 1 skipped, including the two tests that had failed.
    • pytest -n 1 -rs tests/pyathena -k "spark_dataframe or spark_sql": 4 passed, 2 skipped. All three test_spark_dataframe tests read the generated spark_group_by.csv and passed. Two test_spark_sql tests were skipped by pytest-dependency (depends on test_spark_dataframe); the third passed.
  • Not run locally: the full PyAthena suite. AWS CI runs it when the PR is Ready. No SQLAlchemy or Spark path changed, so the changes job skips the SQLAlchemy compliance suites and the Spark tests in CI.

🤖 Generated with Claude Code

laughingman7743 and others added 2 commits September 28, 2026 17:38
The session fixtures were a Jinja DDL template plus hand-made data files,
including a gzipped Hive text row with \002/\003 separators for the complex
columns. tests/pyathena/tables.py now defines each table once: columns with
their Athena and Arrow types, rows as Python values, storage, comment, and
table properties. The session setup generates the DDL and the data files
from it and uploads them; the Spark tests' CSV file is generated as well.

Tables are Parquet except one_row, whose SerDe and delimiters the reflection
tests assert, and many_rows, which tests read without ORDER BY: Athena
returned a Parquet file of 10,000 rows out of file order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread tests/pyathena/tables.py
name: str
columns: tuple[Column, ...]
rows: tuple[tuple[Any, ...], ...] = ()
storage: Literal["parquet", "text"] = "parquet"

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): FINDINGS (1 simplification, fixed)

Base 7b4cceb (#873 head, the stacked base), head 12cfd65. Full diff reviewed (10 files).

  • Finding: Table.storage was a plain str. A misspelled value would silently produce Parquet DDL and data. It is now Literal["parquet", "text"], in 2760077. No behavior change; just lint and mypy tests/pyathena/tables.py pass.
  • DDL:
    • Text tables emit the old ROW FORMAT DELIMITED ... STORED AS TEXTFILE clause verbatim, so the one_row reflection asserts (SerDe, input/output format, field.delim, line.delim, serialization.format) still hold.
    • parquet_with_compression keeps STORED AS PARQUET and TBLPROPERTIES ('parquet.compress'='SNAPPY').
    • CREATE VIEW replaces CREATE OR REPLACE VIEW, which is safe because every session uses a fresh schema.
  • Data:
    • Parquet is written with an explicit schema from Column.arrow_type. The live Athena results match the old files: see the parity check in the PR description.
    • timestamp("ms"), map_ from key/value tuples, decimal128(10,1) and binary all read back as before, and test_complex passes for every cursor.
    • NULLs in integer_na_values/boolean_na_values are Parquet nulls, and the NA tests pass.
  • Row order: many_rows stays text after the observed Parquet reordering. The 3-row NA tables are single-page Parquet files, and their SELECT * order tests pass.
  • Resources: _data_objects is cached per process (ENV.schema is fixed at import), so setup uploads and teardown deletes the same keys. When the upload fails midway, pytest_sessionfinish is skipped, as before this change, and the 1-day lifecycle rule removes the objects.
  • Python 3.10: str | None annotations and zip(strict=True) are 3.10-compatible. pa.Table.from_pylist exists in the pyarrow>=10 floor.

Comment thread tests/pyathena/tables.py
storage="text",
comment="table comment",
),
# A text table: tests read it without ORDER BY and expect the file order,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, callers, AWS operations): FINDINGS (description only, fixed)

Base 7b4cceb, head 2760077.

Claims checked:

  • "Same statements as before": after Give the executemany and partition tests their own tables #873 the template had 7 tables and 2 views, and TABLES/VIEWS have the same 7 and 2.
  • "7 objects per process": before, 6 row files plus test.dat; now 5 data files (partition_table and parquet_with_compression have no rows) plus the Spark CSV plus test.dat.
  • "Byte-identical": checked offline for the one_row, many_rows and spark_group_by.csv bytes.
  • "Out of file order": seen in the first targeted run, where test_pandas_cursor_chunked_vs_regular_same_data got 0..6, 15..30, 7... After the switch to text, the order-sensitive -k set (371 passed) includes both failed tests.
  • Removed-file references: git grep over the repository finds no remaining reference to tests/resources/rows, create_table.sql or the row file names. The license-header config entry is removed, and NOTICE and docs never listed them.
  • CI path filters: nothing under pyathena/(aio/)?sqlalchemy, tests/sqlalchemy, tests/pyathena/(aio/)?sqlalchemy or the Spark paths changed, so CI runs only the PyAthena suite, without Spark tests.

Finding: the TEST section attributed the Spark test_spark_sql skips to xdist without evidence. The skip reason printed by -rs is pytest-dependency's depends on test_spark_dataframe. The description now quotes that reason and notes the Literal-only commit after the tested commit.


def _upload_rows():
@functools.cache
def _data_objects():

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed Codex result): CLEAN

  • Reviewer: Codex CLI 0.157.1 (codex exec -s read-only), model gpt-6-astra, session 01a0e72e-c8d6-7830-88e2-bb7c702d0fd3.
  • Scope: base 7b4cceb, head 2760077. The review ran on a detached snapshot worktree without .env. The prompt left out the PR number, description, commit messages, and self-review findings. After the review, the snapshot was unchanged and still at the reviewed head.
  • Covered surfaces, as reported:
    • the generated DDL and data compared with all removed resources, including the decompressed Hive row;
    • types, values, NULLs, row order, reflection metadata, views, and the consumers of the Spark CSV;
    • S3 keys, upload/delete symmetry, worker isolation, and setup/cleanup failure paths;
    • Python and PyArrow compatibility, workflow dependencies, and remaining references to removed files.
  • Result: "No correctness regressions found by static inspection." This was a static review only: no builds, tests, or network access.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant