Problem
The "Performance tuning options" example in docs/pandas.md passes options that pandas.read_csv does not accept, so two of its calls fail when the results are read:
cursor.execute("SELECT * FROM large_table",
engine="pyarrow",
chunksize=100_000,
use_threads=True)
...
cursor.execute("SELECT * FROM data_table",
engine="c",
chunksize=50_000,
buffer_lines=100_000)
PandasCursor forwards extra keyword arguments to read_csv.
use_threads is not a read_csv parameter. Also, with chunksize set, _get_csv_engine uses the C engine even when engine="pyarrow" is requested.
buffer_lines is not a read_csv parameter either.
Both raise TypeError: read_csv() got an unexpected keyword argument .... The engine="python", low_memory=True example does not raise.
Reproduction
The same arguments passed to pandas directly:
import io
import pandas as pd
pd.read_csv(io.StringIO("a\n1\n2\n"), engine="c", chunksize=2, use_threads=True)
# TypeError: read_csv() got an unexpected keyword argument 'use_threads'
pd.read_csv(io.StringIO("a\n1\n2\n"), engine="c", chunksize=2, buffer_lines=10)
# TypeError: read_csv() got an unexpected keyword argument 'buffer_lines'
Environment
- PyAthena master (
a86a180); the same section is in v3.36.0.
- pandas 3.0.6. Neither option is a
read_csv parameter in pandas 2.x either.
Proposed fix (optional)
Replace the invalid examples with options that read_csv accepts. Also state that engine="pyarrow" is not used together with chunksize, as _get_csv_engine documents. Validate with just docs lint and just docs build.
This was found in an independent review of #885.
Problem
The "Performance tuning options" example in
docs/pandas.mdpasses options thatpandas.read_csvdoes not accept, so two of its calls fail when the results are read:PandasCursorforwards extra keyword arguments toread_csv.use_threadsis not aread_csvparameter. Also, withchunksizeset,_get_csv_engineuses the C engine even whenengine="pyarrow"is requested.buffer_linesis not aread_csvparameter either.Both raise
TypeError: read_csv() got an unexpected keyword argument .... Theengine="python", low_memory=Trueexample does not raise.Reproduction
The same arguments passed to pandas directly:
Environment
a86a180); the same section is in v3.36.0.read_csvparameter in pandas 2.x either.Proposed fix (optional)
Replace the invalid examples with options that
read_csvaccepts. Also state thatengine="pyarrow"is not used together withchunksize, as_get_csv_enginedocuments. Validate withjust docs lintandjust docs build.This was found in an independent review of #885.