Fix analyzedb failure on non-ASCII identifiers - #1938
Draft
tuhaihe wants to merge 1 commit into
Draft
Conversation
analyzedb aborted with UnicodeEncodeError whenever the database held
a schema, table or column whose name is not pure ASCII:
File "analyzedb", line 985, in regclass_schema_tbl
return "to_regclass('%s')" % (pg.escape_string(schema_tbl))
UnicodeEncodeError: 'ascii' codec can't encode characters in
position 13-19: ordinal not in range(128)
The SQL literals were escaped with the module level
pg.escape_string() from PyGreSQL. That function is not bound to a
connection, so it has no client encoding to work with and always
encodes str arguments as ASCII -- see pg_escape_string() in
PyGreSQL's pgmodule.c, which passes pg_encoding_ascii. The
connection method conn.escape_string() uses PQclientEncoding() and
PQescapeStringConn() instead, and handles any encoding. This is
still the case in the latest PyGreSQL, so upgrading the bundled
version would not have helped.
Add dbconn.escapeString(conn, value), which escapes through the
connection, and use it for all five call sites in analyzedb.
get_oid_str() and regclass_schema_tbl() now take the connection so
that they can reach it; every one of their callers already had one
at hand.
Two nearby spots interpolated names into SQL without escaping them
at all, and are now escaped the same way: the schema name given to
"analyzedb -s <schema>", and the schema and table name passed to
GET_LEAF_PARTITIONS_SQL.
The behave scenario meant to cover this ("analyzedb can handle the
table name with special utf-8 characters") only created a TEMP
table, and temp schemas are skipped further down in the flow, so it
had stopped exercising the failing path. Extend it with a permanent
table, a non-ASCII schema and the -s path, and assert that those
tables really appear in the output rather than only checking the
exit code.
The same pattern is still present in gpload.py and in Escape() and
escapeArrayElement() in gppylib/utils.py, which affect gpload, gpsd
and minirepro. Those are left for a follow-up.
Reported-by: vsbace
Assisted-by: Claude Code
See: Issue#1929 <apache#1929>
2 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
analyzedb aborted with UnicodeEncodeError whenever the database held a schema, table or column whose name is not pure ASCII:
File "analyzedb", line 985, in regclass_schema_tbl
return "to_regclass('%s')" % (pg.escape_string(schema_tbl))
UnicodeEncodeError: 'ascii' codec can't encode characters in
position 13-19: ordinal not in range(128)
The SQL literals were escaped with the module level pg.escape_string() from PyGreSQL. That function is not bound to a connection, so it has no client encoding to work with and always encodes str arguments as ASCII -- see pg_escape_string() in PyGreSQL's pgmodule.c, which passes pg_encoding_ascii. The connection method conn.escape_string() uses PQclientEncoding() and PQescapeStringConn() instead, and handles any encoding. This is still the case in the latest PyGreSQL, so upgrading the bundled version would not have helped.
Add dbconn.escapeString(conn, value), which escapes through the connection, and use it for all five call sites in analyzedb. get_oid_str() and regclass_schema_tbl() now take the connection so that they can reach it; every one of their callers already had one at hand.
Two nearby spots interpolated names into SQL without escaping them at all, and are now escaped the same way: the schema name given to "analyzedb -s ", and the schema and table name passed to GET_LEAF_PARTITIONS_SQL.
The behave scenario meant to cover this ("analyzedb can handle the table name with special utf-8 characters") only created a TEMP table, and temp schemas are skipped further down in the flow, so it had stopped exercising the failing path. Extend it with a permanent table, a non-ASCII schema and the -s path, and assert that those tables really appear in the output rather than only checking the exit code.
The same pattern is still present in gpload.py and in Escape() and escapeArrayElement() in gppylib/utils.py, which affect gpload, gpsd and minirepro. Those are left for a follow-up.
Reported-by: vsbace
Assisted-by: Claude Code
See: Issue#1929 #1929
Fixes #ISSUE_Number
What does this PR do?
Type of Change
Breaking Changes
Test Plan
make installcheckmake -C src/test installcheck-cbdb-parallelImpact
Performance:
User-facing changes:
Dependencies:
Checklist
Additional Context
CI Skip Instructions