Skip to content

Add an ARRAY<string> column to one_row_complex through its definition only - #877

Closed
laughingman7743 wants to merge 1 commit into
test/848-derived-expectations-dataframesfrom
test/848-add-column-demo
Closed

laughingman7743 wants to merge 1 commit into
test/848-derived-expectations-dataframesfrom
test/848-add-column-demo

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

WHAT

This PR adds col_array_string ARRAY<string>, with the value ["a", "b"], to ONE_ROW_COMPLEX in tests/pyathena/tables.py. The diff is two lines, both in the table definition: the column and its value. No test or rule changes.

WHY

Part of #848, for #834. Stacked on #876.

#848's validation plan asks to "show that adding a column to a shared table is a one-place change". Before #874, #875, and #876, adding this column meant editing four places:

  1. the Hive text file one_row_complex.gz, which uses \002 separators;
  2. the DDL template;
  3. about 20 hand-written expected rows, schemas, and dtype lists;
  4. the SQLAlchemy column count.

After this change, the tests below cover the new column in every representation.

Test What it now also asserts for col_array_string
TestCursor.test_complex, S3FS ×2, pandas.util.as_pandas ["a", "b"]: the cursors parse Athena's rendering [a, b] into strings
TestSQLAlchemyAthena.test_reflect_select reflected AthenaArray with String items; value ["a", "b"]
PandasCursor CSV "[a, b]" as a pandas string column
PandasCursor UNLOAD ["a", "b"], with dtype object
ArrowCursor CSV (rows, as_arrow, as_polars) "[a, b]", pa.string() / pl.String
ArrowCursor UNLOAD ["a", "b"], with pa.list_(pa.field("array_element", pa.string())) / pl.List(pl.String)
PolarsCursor nothing: arrays are outside the columns these tests select, as before

TEST

Tested commit: bcbf874, then rebased onto the simplified #875/#876 as cbd12b1 (tested again), then onto later #875/#876 heads; current head c3b08bb. The diff is unchanged.

  • just lint passed.
  • Against AWS, from this worktree, pytest -n 4 on tests/pyathena/test_cursor.py, both s3fs cursor test files, pandas/test_util.py, sqlalchemy/test_base.py, and the pandas, arrow, and polars test_cursor.py files, with -k "complex or as_pandas or reflect or table_names or view_names or like or escape or char_length or cast_as_varbinary or filter": 104 passed on bcbf874 and again on cbd12b1. This covers every test that derives one_row_complex expectations, plus the other tests that read one_row_complex or list the schema's tables.
  • Against AWS, pytest -n 4 tests/pyathena/test_glue.py tests/pyathena/test_cursor.py tests/pyathena/aio/test_cursor.py tests/pyathena/sqlalchemy/test_base.py tests/pyathena/aio/sqlalchemy/test_base.py -k "glue or throttl or list_table_metadata or get_table_metadata": 43 passed. These tests compare the schema-wide Glue and Athena metadata, which now include the new column.
  • Not run locally: the full suites. When the PR is Ready, AWS CI runs the PyAthena suite. It skips the SQLAlchemy compliance suites, because no SQLAlchemy path changes here.

🤖 Generated with Claude Code

Comment thread tests/pyathena/tables.py
pa.struct([("a", pa.int32()), ("b", pa.int32())]),
),
Column("col_decimal", "DECIMAL(10,1)", pa.decimal128(10, 1)),
Column("col_array_string", "ARRAY<string>", pa.list_(pa.string())),

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): CLEAN

Base 3202ab5 (#876 head, the stacked base), head bcbf874. The full diff is 2 lines.

  • Generation: Table.data_file writes ["a", "b"] as list<string> and the new round-trip check accepts it. The DDL gains col_array_string ARRAY<string>.
  • Derived tests: all 19 whole-row tests pick the column up through Selection. The PolarsCursor tests leave it out through COMPLEX_FAMILIES.
  • Other readers of the table:
    • The tests that select explicit one_row_complex columns are unaffected.
    • meta.reflect() over the schema (test_get_table_names/test_get_view_names) reflects the new array type.
    • The Glue-vs-Athena and throttled-metadata comparisons list the table with the new column on both sides.
  • AWS results: 104 passed for the derived-expectation, reflection, and one_row_complex column tests, and 43 passed for the Glue, throttling, and table-metadata tests.

Comment thread tests/pyathena/tables.py
[(1, 2), (3, 4)],
{"a": 1, "b": 2},
Decimal("0.1"),
["a", "b"],

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, callers, AWS operations): CLEAN

Base 3202ab5, head bcbf874.

Claims checked against the AWS run on this head

  • Each row of the description's table is a derived expectation that passed on AWS:
    • ["a", "b"] for the Python-object cursors and SQLAlchemy
    • "[a, b]" for the CSV paths
    • pa.list_(pa.field("array_element", pa.string())) / pl.List(pl.String) for Arrow UNLOAD
    • object dtype for pandas UNLOAD
  • "Two lines, both in the table definition": the diff touches only tests/pyathena/tables.py.
  • "Four places before": the gzipped Hive text row, the Jinja DDL, the hand-written expectations in about 20 tests, and len(one_row_complex.c) == 16. All four existed before Define the shared test tables in Python and generate their data #874.

Operational effect: none beyond one more column in one Parquet file. The number of setup statements and uploads is unchanged.

CI scope: no SQLAlchemy or Spark path changes, so on Ready the PyAthena suite runs without Spark.

Comment thread tests/pyathena/tables.py
pa.struct([("a", pa.int32()), ("b", pa.int32())]),
),
Column("col_decimal", "DECIMAL(10,1)", pa.decimal128(10, 1)),
Column("col_array_string", "ARRAY<string>", pa.list_(pa.string())),

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed Codex result): CLEAN

Reviewer: Codex CLI 0.157.1 (codex exec -s read-only), model gpt-6-astra, session 01a0e7b6-1295-7331-87c3-e21bcaa43fd8.

Scope: base 3202ab5, head bcbf874. The review ran on a detached snapshot without .env, and the prompt left out the PR framing. The snapshot was still clean at the head afterwards. This was a static review only.

Covered, as reported:

  • DDL, Parquet generation, and the Arrow round-trip guard.
  • All consumers under tests/, including metadata, Glue, reflection, and schema listings, plus assertions that depend on column count or position.
  • The derived rows, descriptions, pandas dtypes, Arrow schemas, Polars dtypes, and SQLAlchemy array element types.
  • The default, S3FS, pandas, Arrow, and Polars converter and result-set paths, including UNLOAD.

Result: "No broken or newly vacuous assertions found. The new value's expectations match the applicable cursor paths, and ["a", "b"] is losslessly representable by pa.list_(pa.string())."

@laughingman7743
laughingman7743 marked this pull request as ready for review September 28, 2026 11:13
@laughingman7743
laughingman7743 force-pushed the test/848-derived-expectations-dataframes branch from 3202ab5 to 9e350b7 Compare September 28, 2026 14:49
@laughingman7743
laughingman7743 force-pushed the test/848-derived-expectations-dataframes branch from 9e350b7 to 7845405 Compare September 28, 2026 15:01
@laughingman7743
laughingman7743 force-pushed the test/848-derived-expectations-dataframes branch from 7845405 to 68270a4 Compare September 28, 2026 15:07
The column is added only to the table definition. The session creates it,
and the tests that assert whole one_row_complex rows, schemas, dtypes,
descriptions, and reflected types derive its expectations for every cursor
type from the existing rules.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@laughingman7743

Copy link
Copy Markdown
Member Author

Closing without merging, by the maintainer's decision. Deriving the expected results from the table definitions required tests/pyathena/expected.py to re-model how each cursor returns every type: Athena's text rendering, JSON parsing, and the pandas, Arrow, and Polars types. The module grew large enough to need tests of its own, which is more complexity than the one-place column addition is worth. The hand-written expectations stay. See #848 for the summary. The branch is kept for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant