Skip to content

get_athena_type maps Arrow decimal256 to string #888

Description

@laughingman7743

Problem

pyathena.arrow.util.get_athena_type maps an Arrow decimal256 type to string instead of decimal.

if type_.id in [types.Type_DECIMAL128, types.Decimal256Type]:  # 23, 24

types.Decimal256Type is the type class, not the type id types.Type_DECIMAL256, so the comparison never matches a decimal256 field. The field falls through to the final return "string", 2147483647, 0.

to_column_info uses this mapping to build the cursor description from a Parquet schema, as with the unload option of the Arrow and pandas cursors.

Impact is limited in practice. Athena's DECIMAL precision is at most 38, which fits decimal128, so Athena's own UNLOAD output should not contain decimal256. Other Parquet schemas passed to to_column_info can.

Reproduction

import pyarrow as pa
from pyathena.arrow.util import to_column_info

print(to_column_info(pa.schema([
    pa.field("d128", pa.decimal128(10, 2)),
    pa.field("d256", pa.decimal256(40, 2)),
])))
({'Name': 'd128', 'Type': 'decimal', 'Precision': 10, 'Scale': 2, 'Nullable': 'NULLABLE'},
 {'Name': 'd256', 'Type': 'string', 'Precision': 2147483647, 'Scale': 0, 'Nullable': 'NULLABLE'})

Environment

  • PyAthena master (a86a180); the same code is in v3.36.0.
  • pyarrow 25.0.1, Python 3.13.

Proposed fix (optional)

Compare with types.Type_DECIMAL256, and add a unit test for decimal256 next to the existing get_athena_type / to_column_info tests. No AWS resources are needed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions