Skip to content

Redesign the PyAthena test fixtures: define once, generate data, set up once per job #848

Description

@laughingman7743

Use case

The PyAthena suite's shared fixtures have barely changed since 2022 (#295). They are costly to set up, and they are hard to change.

Setup cost

  • tests/__init__.py gives every process its own random schema name (pyathena_test_<random>).
  • pytest_sessionstart in tests/pyathena/conftest.py runs in every process. It uploads tests/resources/rows/* to S3, creates the database, and runs create_table.sql.jinja2: 34 statements, namely 16 DROP TABLE IF EXISTS (redundant in a fresh schema), 16 CREATE EXTERNAL TABLE, and 2 CREATE OR REPLACE VIEW.
  • With pytest -n 8, one job sets up the same fixtures 9 times: 8 workers plus the pytest-xdist controller, which runs no tests. A full run on 5 Python versions creates 45 copies. In the sample on Reduce the AWS cost of integration tests without reducing real-service coverage #834, CREATE TABLE/DROP TABLE statements were 15.4% of all CI queries.

Maintainability

  • The data is files in tests/resources/rows/, prepared by hand in advance. The maintainer has confirmed that this approach does not need to be kept:
    • one_row_complex.gz is a single gzipped row in Hive text format, with \002/\003 separators for arrays, maps, and structs. It cannot practically be edited by hand.
    • many_rows.tsv has 10,000 lines.
  • one_row_complex (16 columns) is referenced 83 times in 9 files. The cursor tests (default, pandas, arrow, polars, s3fs) spell out the full row or DataFrame for each cursor type, and the SQLAlchemy reflection tests assert its column list. Adding a column means editing the Hive text file and every one of those expectations.
  • 9 of the 16 tables are execute_many* tables: mutable, one per cursor type, created for every session whether used or not.
  • tests/resources/queries/insert_into_table.sql.jinja2 is not referenced anywhere.

Proposed change

Redesign the shared fixtures so that they are defined once, generated, and set up once per job.

  1. Single definition in Python.
    • Describe each shared table in code: name, columns with Athena types, storage format, partitions, and rows as Python values.
    • Generate the DDL and the data files from this definition, and upload them. No queries are needed beyond CREATE EXTERNAL TABLE/CREATE VIEW.
    • The data format is free to choose; TSV does not need to be kept. Parquet written with pyarrow (already a dev dependency) is a good default: it represents binary, decimal, microsecond timestamps, and nested array/map/struct values natively, with no delimiter escaping.
    • Keep text-format tables only where a test depends on text-format behavior, such as NA handling.
  2. Expectations derived from the definition.
    • Tests that read a whole row get their expected values from the definition, with per-cursor conversion helpers (Python objects, pandas, Arrow, Polars).
    • Other tests select explicit columns instead of SELECT *.
    • Adding a column or a type-coverage table then means adding one entry.
  3. Set up once per job.
    • Skip the setup in the xdist controller and drop the redundant DROP TABLE IF EXISTS statements (a small first step).
    • Then create the read-only fixtures once, in one schema shared by all workers. For example, the controller creates it and passes the name through pytest-xdist workerinput.
    • Keep cleanup working after interrupted runs (the scheduled database sweep covers leftovers).
  4. Mutable tables per test. Replace the session-wide execute_many* tables with per-test tables, like the existing executemany_table fixture, so the shared schema stays read-only.
  5. Remove the unused insert_into_table.sql.jinja2.

This can land in steps. Step 3's first part is independent and cheap; the definition and the expectation helpers can move one table at a time, starting with one_row_complex.

Validation plan (if implementing)

  • just test pyathena against AWS passes with -n 8 and with -n 1, with the same number of tests and no weakened assertions. Compare the collected test IDs and the pass/skip counts before and after.
  • Count the setup statements and S3 uploads per job before and after.
  • Show that adding a column to a shared table is a one-place change, by adding one in the PR that introduces the definition, or in a follow-up.
  • Confirm that the schema and the S3 Tables namespace are cleaned up after normal and interrupted runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions