Skip to content

Share the pure logic duplicated between sync and asyncio cursors #880

Description

@laughingman7743

Use case

Most asyncio cursor and result-set methods in pyathena/aio/ are near-copies of their sync counterparts.
Their bodies differ only in await, retry_api_call versus async_retry_api_call, time.sleep versus asyncio.sleep, and KeyboardInterrupt versus asyncio.CancelledError.
Each copy carries the same request building, response parsing, state checks, and log messages, so a fix in one version can miss the other.
One drift already happened: the async Spark cursor logged Failed to terminate session. while the sync one logged the session ID. #879 aligns that message.

Approximate duplicated method bodies, counted on the #879 branch (docstrings excluded; the docstrings are duplicated too):

Area Sync / async files Duplicated pairs Lines
Cursor base common.py, result_set.py (WithFetch) / aio/common.py 17, plus 10 verbatim WithFetch / WithAsyncFetch members ~380
Spark common.py, spark/common.py, spark/cursor.py / aio/spark/cursor.py 15 ~130
Result set result_set.py / aio/result_set.py 7 ~65
Arrow, pandas, Polars, S3FS cursors {arrow,pandas,polars,s3fs}/cursor.py / aio/{...}/cursor.py ~16, plus fetchone / fetchmany / fetchall repeated in each of the 5 aio cursors ~390

The aio format cursors reuse the sync result sets through asyncio.to_thread, so the format result sets themselves are not duplicated.

Proposed change

Move the pure (non-I/O) parts into single-underscore helpers on the existing base classes, following the existing _build_* request builders.
Keep separate sync and async methods for the I/O calls themselves; do not add a sans-I/O layer or a code generator.
Candidates, largest first:

  • _find_previous_query_id: the cache-window setup and the filter / sort / expiry / match loop are pure; only _list_query_executions is I/O.
  • The execute prologue (ExecuteOptions.resolve, _prepare_query / _prepare_unload, _build_start_query_execution_request), repeated in _execute and in the 10 format cursors' execute.
  • The result-set kwargs built in each format cursor's execute.
  • The verbatim WithFetch / WithAsyncFetch members (arraysize, rownumber, rowcount, close, and so on), which could live on WithResultSet.
  • The aio format cursors' fetchone / fetchmany / fetchall, which differ only in a cast.
  • Response parsing in list_databases, list_table_metadata, and _batch_get_query_execution.

Keep the signatures of the single-underscore methods that downstream code calls (_execute, _poll, _cancel, _get_query_execution, _find_previous_query_id; see #734).
No behavior change. Split the work into one PR per area in the table.

Validation plan (if implementing)

  • just lint.
  • For each area, the matching sync and aio PyAthena tests on AWS (cursor, aio cursor, Spark, and the format cursors), plus just test sqla and just test sqla-async when the format cursors change.
  • Re-count the duplicated bodies per area to show the reduction.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions