Skip to content

Fix analyzedb failure on non-ASCII identifiers - #1938

Draft
tuhaihe wants to merge 1 commit into
apache:mainfrom
tuhaihe:fix-1929-analyzedb-non-ascii
Draft

Fix analyzedb failure on non-ASCII identifiers#1938
tuhaihe wants to merge 1 commit into
apache:mainfrom
tuhaihe:fix-1929-analyzedb-non-ascii

Conversation

@tuhaihe

@tuhaihe tuhaihe commented Aug 28, 2026

Copy link
Copy Markdown
Member

analyzedb aborted with UnicodeEncodeError whenever the database held a schema, table or column whose name is not pure ASCII:

File "analyzedb", line 985, in regclass_schema_tbl
return "to_regclass('%s')" % (pg.escape_string(schema_tbl))
UnicodeEncodeError: 'ascii' codec can't encode characters in
position 13-19: ordinal not in range(128)

The SQL literals were escaped with the module level pg.escape_string() from PyGreSQL. That function is not bound to a connection, so it has no client encoding to work with and always encodes str arguments as ASCII -- see pg_escape_string() in PyGreSQL's pgmodule.c, which passes pg_encoding_ascii. The connection method conn.escape_string() uses PQclientEncoding() and PQescapeStringConn() instead, and handles any encoding. This is still the case in the latest PyGreSQL, so upgrading the bundled version would not have helped.

Add dbconn.escapeString(conn, value), which escapes through the connection, and use it for all five call sites in analyzedb. get_oid_str() and regclass_schema_tbl() now take the connection so that they can reach it; every one of their callers already had one at hand.

Two nearby spots interpolated names into SQL without escaping them at all, and are now escaped the same way: the schema name given to "analyzedb -s ", and the schema and table name passed to GET_LEAF_PARTITIONS_SQL.

The behave scenario meant to cover this ("analyzedb can handle the table name with special utf-8 characters") only created a TEMP table, and temp schemas are skipped further down in the flow, so it had stopped exercising the failing path. Extend it with a permanent table, a non-ASCII schema and the -s path, and assert that those tables really appear in the output rather than only checking the exit code.

The same pattern is still present in gpload.py and in Escape() and escapeArrayElement() in gppylib/utils.py, which affect gpload, gpsd and minirepro. Those are left for a follow-up.

Reported-by: vsbace
Assisted-by: Claude Code
See: Issue#1929 #1929

Fixes #ISSUE_Number

What does this PR do?

Type of Change

  • Bug fix (non-breaking change)
  • New feature (non-breaking change)
  • Breaking change (fix or feature with breaking changes)
  • Documentation update

Breaking Changes

Test Plan

  • Unit tests added/updated
  • Integration tests added/updated
  • Passed make installcheck
  • Passed make -C src/test installcheck-cbdb-parallel

Impact

Performance:

User-facing changes:

Dependencies:

Checklist

Additional Context

CI Skip Instructions


analyzedb aborted with UnicodeEncodeError whenever the database held
a schema, table or column whose name is not pure ASCII:

  File "analyzedb", line 985, in regclass_schema_tbl
    return "to_regclass('%s')" % (pg.escape_string(schema_tbl))
  UnicodeEncodeError: 'ascii' codec can't encode characters in
  position 13-19: ordinal not in range(128)

The SQL literals were escaped with the module level
pg.escape_string() from PyGreSQL. That function is not bound to a
connection, so it has no client encoding to work with and always
encodes str arguments as ASCII -- see pg_escape_string() in
PyGreSQL's pgmodule.c, which passes pg_encoding_ascii. The
connection method conn.escape_string() uses PQclientEncoding() and
PQescapeStringConn() instead, and handles any encoding. This is
still the case in the latest PyGreSQL, so upgrading the bundled
version would not have helped.

Add dbconn.escapeString(conn, value), which escapes through the
connection, and use it for all five call sites in analyzedb.
get_oid_str() and regclass_schema_tbl() now take the connection so
that they can reach it; every one of their callers already had one
at hand.

Two nearby spots interpolated names into SQL without escaping them
at all, and are now escaped the same way: the schema name given to
"analyzedb -s <schema>", and the schema and table name passed to
GET_LEAF_PARTITIONS_SQL.

The behave scenario meant to cover this ("analyzedb can handle the
table name with special utf-8 characters") only created a TEMP
table, and temp schemas are skipped further down in the flow, so it
had stopped exercising the failing path. Extend it with a permanent
table, a non-ASCII schema and the -s path, and assert that those
tables really appear in the output rather than only checking the
exit code.

The same pattern is still present in gpload.py and in Escape() and
escapeArrayElement() in gppylib/utils.py, which affect gpload, gpsd
and minirepro. Those are left for a follow-up.

Reported-by: vsbace
Assisted-by: Claude Code
See: Issue#1929 <apache#1929>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant