You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
pyathena.pandas.util.to_sql() does not find an existing table whose name was given with uppercase letters, so if_exists="fail" does not fail and data is appended twice.
It passes schema and name through unchanged. Athena stores names in lowercase and reports them in lowercase in information_schema, so name="MyTable" never matches the existing mytable.
Reproduced against live Athena on master (3c69daf):
to_sql(df, "PyAthenaToSqlCase…", conn, location, schema="default", if_exists="fail") # creates the tableto_sql(df, "PyAthenaToSqlCase…", conn, location, schema="default", if_exists="fail") # no error# SELECT count(*) -> 2 (one row written twice)
if_exists="fail" writes Parquet into the existing table's location and runs CREATE EXTERNAL TABLE IF NOT EXISTS, which is a no-op, so the rows are appended instead of raising.
if_exists="replace" follows the same path: the existing table is not dropped and its S3 objects are not deleted, so it appends too. This follows from the code and was not reproduced separately.
The existing test_to_sql uses a lowercase generated name, so it cannot see this.
Same query, same defect class
The check runs through conn.cursor() without disabling query result reuse. A connection with result_reuse_enable=True can get a cached "no rows" answer from before the table existed, and the existence check fails the same way. Fix metadata reflection throttling and reuse listed metadata #777 turned result reuse off for the dialect's information_schema fallback for this reason.
schema and name are interpolated into string literals without escaping '.
Proposed fix
Lowercase and quote-escape both identifiers in the existence check, the way _columns_from_information_schema does in the SQLAlchemy dialect. Execute the check with result_reuse_enable=False. Add a regression test that runs to_sql twice with a mixed-case name: once expecting if_exists="fail" to raise, and once expecting if_exists="replace" to leave one copy of the rows.
Problem
pyathena.pandas.util.to_sql()does not find an existing table whose name was given with uppercase letters, soif_exists="fail"does not fail and data is appended twice.to_sqlchecks existence with:It passes
schemaandnamethrough unchanged. Athena stores names in lowercase and reports them in lowercase ininformation_schema, soname="MyTable"never matches the existingmytable.Reproduced against live Athena on
master(3c69daf):if_exists="fail"writes Parquet into the existing table's location and runsCREATE EXTERNAL TABLE IF NOT EXISTS, which is a no-op, so the rows are appended instead of raising.if_exists="replace"follows the same path: the existing table is not dropped and its S3 objects are not deleted, so it appends too. This follows from the code and was not reproduced separately.The existing
test_to_sqluses a lowercase generated name, so it cannot see this.Same query, same defect class
conn.cursor()without disabling query result reuse. A connection withresult_reuse_enable=Truecan get a cached "no rows" answer from before the table existed, and the existence check fails the same way. Fix metadata reflection throttling and reuse listed metadata #777 turned result reuse off for the dialect'sinformation_schemafallback for this reason.schemaandnameare interpolated into string literals without escaping'.Proposed fix
Lowercase and quote-escape both identifiers in the existence check, the way
_columns_from_information_schemadoes in the SQLAlchemy dialect. Execute the check withresult_reuse_enable=False. Add a regression test that runsto_sqltwice with a mixed-case name: once expectingif_exists="fail"to raise, and once expectingif_exists="replace"to leave one copy of the rows.Out of scope
AwsDataCatalog,information_schemafilters by Lake Formation rather than failing, so a table the caller cannot see is not found either. This is the same limit Read the dialect's own queries through an API cursor and resolve federated absence from information_schema #798 documented for the dialect.DROP TABLEandALTER TABLEstatements quote identifiers with backticks. Athena resolves those case-insensitively, so they are unaffected.