Use case
Most asyncio cursor and result-set methods in pyathena/aio/ are near-copies of their sync counterparts.
Their bodies differ only in await, retry_api_call versus async_retry_api_call, time.sleep versus asyncio.sleep, and KeyboardInterrupt versus asyncio.CancelledError.
Each copy carries the same request building, response parsing, state checks, and log messages, so a fix in one version can miss the other.
One drift already happened: the async Spark cursor logged Failed to terminate session. while the sync one logged the session ID. #879 aligns that message.
Approximate duplicated method bodies, counted on the #879 branch (docstrings excluded; the docstrings are duplicated too):
| Area |
Sync / async files |
Duplicated pairs |
Lines |
| Cursor base |
common.py, result_set.py (WithFetch) / aio/common.py |
17, plus 10 verbatim WithFetch / WithAsyncFetch members |
~380 |
| Spark |
common.py, spark/common.py, spark/cursor.py / aio/spark/cursor.py |
15 |
~130 |
| Result set |
result_set.py / aio/result_set.py |
7 |
~65 |
| Arrow, pandas, Polars, S3FS cursors |
{arrow,pandas,polars,s3fs}/cursor.py / aio/{...}/cursor.py |
~16, plus fetchone / fetchmany / fetchall repeated in each of the 5 aio cursors |
~390 |
The aio format cursors reuse the sync result sets through asyncio.to_thread, so the format result sets themselves are not duplicated.
Proposed change
Move the pure (non-I/O) parts into single-underscore helpers on the existing base classes, following the existing _build_* request builders.
Keep separate sync and async methods for the I/O calls themselves; do not add a sans-I/O layer or a code generator.
Candidates, largest first:
_find_previous_query_id: the cache-window setup and the filter / sort / expiry / match loop are pure; only _list_query_executions is I/O.
- The execute prologue (
ExecuteOptions.resolve, _prepare_query / _prepare_unload, _build_start_query_execution_request), repeated in _execute and in the 10 format cursors' execute.
- The result-set kwargs built in each format cursor's
execute.
- The verbatim
WithFetch / WithAsyncFetch members (arraysize, rownumber, rowcount, close, and so on), which could live on WithResultSet.
- The aio format cursors'
fetchone / fetchmany / fetchall, which differ only in a cast.
- Response parsing in
list_databases, list_table_metadata, and _batch_get_query_execution.
Keep the signatures of the single-underscore methods that downstream code calls (_execute, _poll, _cancel, _get_query_execution, _find_previous_query_id; see #734).
No behavior change. Split the work into one PR per area in the table.
Validation plan (if implementing)
just lint.
- For each area, the matching sync and aio PyAthena tests on AWS (cursor, aio cursor, Spark, and the format cursors), plus
just test sqla and just test sqla-async when the format cursors change.
- Re-count the duplicated bodies per area to show the reduction.
Use case
Most asyncio cursor and result-set methods in
pyathena/aio/are near-copies of their sync counterparts.Their bodies differ only in
await,retry_api_callversusasync_retry_api_call,time.sleepversusasyncio.sleep, andKeyboardInterruptversusasyncio.CancelledError.Each copy carries the same request building, response parsing, state checks, and log messages, so a fix in one version can miss the other.
One drift already happened: the async Spark cursor logged
Failed to terminate session.while the sync one logged the session ID. #879 aligns that message.Approximate duplicated method bodies, counted on the #879 branch (docstrings excluded; the docstrings are duplicated too):
common.py,result_set.py(WithFetch) /aio/common.pyWithFetch/WithAsyncFetchmemberscommon.py,spark/common.py,spark/cursor.py/aio/spark/cursor.pyresult_set.py/aio/result_set.py{arrow,pandas,polars,s3fs}/cursor.py/aio/{...}/cursor.pyfetchone/fetchmany/fetchallrepeated in each of the 5 aio cursorsThe aio format cursors reuse the sync result sets through
asyncio.to_thread, so the format result sets themselves are not duplicated.Proposed change
Move the pure (non-I/O) parts into single-underscore helpers on the existing base classes, following the existing
_build_*request builders.Keep separate sync and async methods for the I/O calls themselves; do not add a sans-I/O layer or a code generator.
Candidates, largest first:
_find_previous_query_id: the cache-window setup and the filter / sort / expiry / match loop are pure; only_list_query_executionsis I/O.ExecuteOptions.resolve,_prepare_query/_prepare_unload,_build_start_query_execution_request), repeated in_executeand in the 10 format cursors'execute.execute.WithFetch/WithAsyncFetchmembers (arraysize,rownumber,rowcount,close, and so on), which could live onWithResultSet.fetchone/fetchmany/fetchall, which differ only in acast.list_databases,list_table_metadata, and_batch_get_query_execution.Keep the signatures of the single-underscore methods that downstream code calls (
_execute,_poll,_cancel,_get_query_execution,_find_previous_query_id; see #734).No behavior change. Split the work into one PR per area in the table.
Validation plan (if implementing)
just lint.just test sqlaandjust test sqla-asyncwhen the format cursors change.