Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions docs/sqlalchemy.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,12 +100,16 @@ A later listing preserves metadata already fetched for a table.
Table-metadata lookups propagate throttling and permission errors rather than reporting missing tables.
`has_table()` propagates permission failures, including access denied by Lake Formation, instead of returning or caching `False`.
For failed metadata requests, only recognized `EntityNotFoundException` responses establish absence; unrecognized errors are propagated rather than guessed to mean a missing table.
Column reflection and `has_table()` do not retry a throttled table-metadata request; when Athena reports throttling, they read `information_schema.columns` instead, executed without query result reuse, and log a warning.
Other error codes listed in the connection's `RetryConfig.exceptions` are still retried on that path, except `MetadataException` itself, which carries wrapped throttling; list the wrapped Glue codes instead.
A `retry_config` in `cursor_kwargs` replaces that policy entirely, including its throttling retries, which then run before the fallback.
Column reflection and `has_table()` do not retry a table-metadata request that `information_schema` can answer; they read `information_schema.columns` instead, executed without query result reuse, and log a warning.
That covers a throttled request in any catalog.
It also covers a `MetadataException` carrying no recognized Glue error envelope, but only outside `AwsDataCatalog`: a federated catalog reports a missing table in its connector's own words, so absence is decided by the query against that catalog rather than by an unrecognized message.
In `AwsDataCatalog` an unrecognized `MetadataException` still propagates, because Glue does state missing tables and permission failures in a recognized envelope, and `information_schema` filters by Lake Formation instead of failing, so reading it there would report a table the caller cannot see as absent.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two — claims and operational behavior. Base 15325e172714e2e809dd5c597e8babb2c51a6fac, head 205c9b14d4e152b39c810609455852f22d38512d. Claims audited: this docs paragraph, the docstrings of _internal_cursor and _answerable_from_information_schema, the commit messages, and the PR body.

Result: FINDINGS, corrected.

  1. The docs still said any MetadataException without an envelope goes to the fallback, after the code had been narrowed to non-Glue catalogs. Corrected in 8ecd6bf. The paragraph now states the catalog boundary and the Lake Formation reason for it.
  2. The claim "only outside AwsDataCatalog" was not yet checked against S3 Tables, which is outside AwsDataCatalog but is Glue-federated. Measured live: a missing S3 Tables table returns Entity Not Found (Service: AmazonDataCatalog; ...; Error Code: EntityNotFoundException; ...), so _lookup_table maps it before the fallback. S3 Tables behavior is unchanged. The PR body says so.
  3. get_view_definition against a missing view was measured live on rest and pandas: NoSuchTableError on both. The stub test for that path cites this.

Operator view: the fallback adds one information_schema query per unrecognized federated failure, instead of that failure's retry ladder. Internal queries no longer UNLOAD, which removes an S3 write and adds no query.

Deferred and stated in the PR: extra_info partition marking and case-preserving connectors on federated catalogs remain unverified. The raw pyathena.error.* from get_view_definition is not SQLAlchemy-wrapped. A cursor_kwargs converter still reaches the internal cursor.

Other error codes listed in the connection's `RetryConfig.exceptions` are still retried on that path, except those; list the wrapped Glue codes instead of `MetadataException`.
A `retry_config` in `cursor_kwargs` replaces that policy entirely, including those retries, which then run before the fallback.
The fallback maps unbounded `varchar` to SQLAlchemy `String`, matching Hive `STRING` reflection from the metadata API, and preserves explicit `VARCHAR(n)` and `CHAR(n)` lengths.
Partition columns are marked from the `extra_info` column.
This fallback does not populate the table-metadata cache, and table comments and table options still come from the metadata API with the configured retries, so they propagate the throttling error.
This fallback does not populate the table-metadata cache, and table comments and table options still come from the metadata API with the configured retries, so they propagate the error.
The dialect runs its own queries — this fallback and `get_view_definition()` — through the API cursor, whatever `cursor_class` or `unload` setting the connection carries, because it parses those result rows itself.
Athena applies its metadata API rate limits per account, and they are not listed in Service Quotas.
PyAthena's API retries use exponential backoff with uniform jitter; `RetryConfig` documents the default attempt count and waits.
PyAthena recognizes Glue error codes in Athena's `MetadataException` service-error envelope and applies `RetryConfig.exceptions` to those codes.
Expand Down
11 changes: 9 additions & 2 deletions pyathena/aio/sqlalchemy/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,8 @@

import pyathena
from pyathena.aio.connection import AioConnection
from pyathena.aio.cursor import AioCursor
from pyathena.cursor import Cursor
from pyathena.error import (
DatabaseError,
DataError,
Expand All @@ -30,6 +32,9 @@
from sqlalchemy import URL


_ASYNC_CURSOR_CLASSES: dict[Any, Any] = {Cursor: AioCursor}


class AsyncAdapt_pyathena_cursor:
"""Wraps any async PyAthena cursor with a sync DBAPI interface.

Expand Down Expand Up @@ -147,8 +152,10 @@ def cursor_kwargs(self) -> dict[str, Any]:
def retry_config(self) -> RetryConfig:
return self._connection.retry_config # type: ignore[no-any-return]

def cursor(self, **kwargs: Any) -> AsyncAdapt_pyathena_cursor:
raw_cursor = self._connection.cursor(**kwargs)
def cursor(self, cursor: Any = None, **kwargs: Any) -> AsyncAdapt_pyathena_cursor:
# The shared dialect names a cursor class in its synchronous form; this
# connection can only drive the async counterpart.
raw_cursor = self._connection.cursor(_ASYNC_CURSOR_CLASSES.get(cursor, cursor), **kwargs)
return AsyncAdapt_pyathena_cursor(raw_cursor)

def close(self) -> None:
Expand Down
117 changes: 88 additions & 29 deletions pyathena/sqlalchemy/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
cast,
)

from sqlalchemy import exc, schema, text, types, util
from sqlalchemy import exc, schema, types, util
from sqlalchemy.engine import Engine, reflection
from sqlalchemy.engine.default import DefaultDialect
from sqlalchemy.engine.interfaces import ExecutionContext
Expand All @@ -23,6 +23,7 @@
)

import pyathena
from pyathena.cursor import Cursor
from pyathena.sqlalchemy.compiler import (
AthenaDDLCompiler,
AthenaStatementCompiler,
Expand Down Expand Up @@ -195,6 +196,12 @@ class AthenaDialect(DefaultDialect):

_connect_options: dict[str, Any] = {} # type: ignore[override] # noqa: RUF012
_pattern_column_type: Pattern[str] = re.compile(r"^([a-zA-Z]+)(?:$|[\(|<](.+)[\)|>]$)")
# Metadata failures that information_schema answers better than a retry.
# Throttling, because one query costs less than the retry ladder, and a
# MetadataException that survived unwrapping, because a federated catalog
# reports a missing table in its connector's words rather than in Glue's
# EntityNotFoundException envelope.
_FALLBACK_ERROR_CODES: tuple[str, ...] = (*THROTTLING_ERROR_CODES, "MetadataException")

def __init__(self, json_deserializer=None, json_serializer=None, **kwargs):
DefaultDialect.__init__(self, **kwargs)
Expand Down Expand Up @@ -347,13 +354,13 @@ def _get_columns(self, connection, table_name: str, schema: str | None = None, *
metadata = info_cache.get(metadata_key)
if metadata is not None:
return self._columns_from_metadata(metadata)
# A throttled metadata request switches to information_schema at once
# instead of waiting out the retry policy; the query answers existence
# and columns, while table comments and options still need the API.
# Other retryable codes keep the connection's policy. Connection.cursor()
# applies cursor_kwargs last, so a retry_config given there still runs
# its own throttling retries before the fallback.
retry_config = self._without_throttling_retries(
# A metadata request the fallback can answer switches to
# information_schema at once instead of waiting out the retry policy; the
# query answers existence and columns, while table comments and options
# still need the API. Other retryable codes keep the connection's policy.
# Connection.cursor() applies cursor_kwargs last, so a retry_config given
# there still runs its own retries before the fallback.
retry_config = self._without_fallback_retries(
raw_connection.retry_config # type: ignore[union-attr]
)
with raw_connection.driver_connection.cursor( # type: ignore[union-attr]
Expand All @@ -362,13 +369,11 @@ def _get_columns(self, connection, table_name: str, schema: str | None = None, *
try:
metadata = self._lookup_table(cursor, schema, name, table_name)
except pyathena.error.OperationalError as e:
if (
_get_error_code(e.__cause__ or e, unwrap_metadata=True)
not in THROTTLING_ERROR_CODES
):
code = _get_error_code(e.__cause__ or e, unwrap_metadata=True)
if not self._is_fallback_error(code, catalog):
raise
_logger.warning(
f"Table metadata request for {table_name} was throttled; "
f"Table metadata request for {table_name} failed with {code}; "
"reflecting columns from information_schema."
)
columns = self._columns_from_information_schema(raw_connection, schema, name)
Expand All @@ -380,21 +385,61 @@ def _get_columns(self, connection, table_name: str, schema: str | None = None, *
return self._columns_from_metadata(metadata)

@staticmethod
def _without_throttling_retries(retry_config: RetryConfig) -> RetryConfig:
"""Copy a policy without the codes that carry throttling.
def _is_fallback_error(code: str | None, catalog: str | None) -> bool:
"""Whether the information_schema fallback answers this failed request.

Athena wraps Glue throttling in ``MetadataException``, so that code is
dropped as well; specific wrapped codes stay retryable.
The codes are those in ``_FALLBACK_ERROR_CODES``; the catalog decides
whether an unrecognized ``MetadataException`` qualifies.

Throttling always: one query costs less than the retry ladder.

A ``MetadataException`` that survived unwrapping only outside the Glue
Data Catalog. Glue states missing tables and permission failures in an
envelope this client recognizes, so an unrecognized one there has an
unknown cause, and answering it from ``information_schema`` would report
a table the caller merely cannot see as absent. A federated catalog has
no such envelope: it reports a missing table in its connector's own
words, which cannot be recognized at all.
"""
if code in THROTTLING_ERROR_CODES:
return True
return code == "MetadataException" and (catalog or "").lower() != "awsdatacatalog"

@classmethod
def _without_fallback_retries(cls, retry_config: RetryConfig) -> RetryConfig:
"""Copy a policy without the codes the information_schema fallback answers.

Retrying those spends the policy's whole budget on a question one query
settles; specific wrapped Glue codes stay retryable.
"""
excluded = (*THROTTLING_ERROR_CODES, "MetadataException")
return RetryConfig(
exceptions=[c for c in retry_config.exceptions if c not in excluded],
exceptions=[c for c in retry_config.exceptions if c not in cls._FALLBACK_ERROR_CODES],
attempt=retry_config.attempt,
multiplier=retry_config.multiplier,
max_delay=retry_config.max_delay,
exponential_base=retry_config.exponential_base,
)

@staticmethod
def _internal_cursor(raw_connection: PoolProxiedConnection) -> Any:
"""Open an API cursor for the queries this dialect parses itself.

Reflection reads these rows directly, so they must not arrive in the
result format chosen for user queries: a DataFrame cursor reports a NULL
or blank value as NaN, as an empty string, or as a dropped row depending
on its backend and on UNLOAD.

The converter is pinned too: a connection-level one chosen for a
DataFrame cursor would otherwise be applied to this one.

The async connection adapter maps ``Cursor`` to its own counterpart and
returns its wrapper, so this is typed by the interface used here rather
than by the class requested.
"""
return raw_connection.driver_connection.cursor( # type: ignore[union-attr]
Cursor, converter=Cursor.get_default_converter()
)

def _column(self, name: str | None, type_: str, comment: str | None, partition: bool | None):
return {
"name": name,
Expand All @@ -420,17 +465,18 @@ def _columns_from_information_schema(
# The answer must reflect the catalog now, so query result reuse is off.
schema = str(schema).lower().replace("'", "''")
table_name = table_name.lower().replace("'", "''")
with raw_connection.driver_connection.cursor() as cursor: # type: ignore[union-attr]
with self._internal_cursor(raw_connection) as cursor:
cursor.execute(
"SELECT ordinal_position, column_name, data_type, comment, extra_info "
"FROM information_schema.columns "
f"WHERE table_schema = '{schema}' AND table_name = '{table_name}'",
result_reuse_enable=False,
)
rows = cursor.fetchall()
# Sort here: the query has no ORDER BY and UNLOAD-backed cursors do not
# preserve result order. A CSV-backed pandas cursor reads a missing
# comment as NaN, which _column() cannot recognize as empty.
# Sort here: the query has no ORDER BY, so its result order is Athena's.
# The comment is still normalized at this boundary: a converter given in
# cursor_kwargs is applied after the one _internal_cursor() pins, and one
# written for a DataFrame cursor reports a missing value as NaN.
return [
self._column(
column_name,
Expand Down Expand Up @@ -523,12 +569,25 @@ def get_view_definition(
raw_connection = self._raw_connection(connection)
schema = schema if schema else self._cursor_option(raw_connection, "schema_name")
query = f"""SHOW CREATE VIEW "{schema}"."{view_name}";"""
try:
res = connection.scalars(text(query))
except exc.OperationalError as e:
raise exc.NoSuchTableError(f"{schema}.{view_name}") from e
else:
return "\n".join(res)
with self._internal_cursor(raw_connection) as cursor:
try:
cursor.execute(query)
except pyathena.error.OperationalError as e:
# Athena runs SHOW CREATE VIEW for a missing view and fails the
# query, which the cursor reports without an underlying API
# error. execute() also fetches the first result page, and a
# failed API call there carries its error as the cause: that is
# a failed read of a view that exists, not a missing one. Any
# query that ends without success is still read as absence, as
# before; its state is not kept and its error codes do not
# single out a missing view.
if e.__cause__ is not None:
raise
raise exc.NoSuchTableError(f"{schema}.{view_name}") from e
rows = cursor.fetchall()

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review — relayed result. Reviewer: Codex, a different model from the author, through the repository's Codex rescue agent. Read-only static inspection, with no edits, tests or GitHub access. Reviewed d7e1dcd..5f87ed3; the rebased equivalent of that head is 8ecd6bf. Surfaces the reviewer read: all changed files, plus util.py, connection.py, common.py, cursor.py, converter.py and result_set.py, the async connection/cursor/adapter paths, and the pandas/arrow/polars cursor and result-set paths.

Result: FINDINGS, 3 items, all repaired in 205c9b1.

  1. P2. fetchall() sat inside the OperationalError -> NoSuchTableError handler. A GetQueryResults page that exhausted its retries (result_set.py:370) reported an existing view as absent. This was a regression introduced by this branch: before it, iteration was outside the handler. Now only execute() is inside the handler.
  2. P2. Connection.cursor() applies cursor_kwargs after the explicit arguments, so a cursor_kwargs["converter"] still overrides the pinned one. Dropping the isinstance(comment, str) guard had therefore reopened the NaN-comment path. The guard is restored at line 482, and the stub test has its NaN row back.
  3. P2. The blank-line test could skip on the defect itself. Losing every blank row left CREATE VIEW and UNION ALL, so both assertions passed and the skip blamed Athena's formatting. The test now compares the full definition against rows fetched through an API cursor (line 949), and skips only when those rows have no blank line.

A follow-up on 8ecd6bf..205c9b1 has been requested and will be recorded in reply.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-ups on the repairs (same Codex reviewer, read-only, static inspection).

Follow-up 1, 8ecd6bf..205c9b1: item 2 (comment guard) CONFIRMED-RESOLVED. Two new FINDINGS, both repaired in 185f582:

  • Cursor.execute() prefetches the first result page (result_set.py _pre_fetch), so narrowing the handler to execute() still caught a failed first GetQueryResults as a missing view. Measured live: a missing view is a query Athena runs and fails (View not found or not a valid presto view: ...), raised as OperationalError with __cause__ is None. An API failure is raised from its ClientError. Only the cause-less form maps to NoSuchTableError now. My earlier missing-view stub had modelled a rejected StartQueryExecution, which actually raises DatabaseError and never reaches this handler. It is replaced, and a failed-first-page test is added.
  • The live test's baseline is read through the same API-cursor path as the fix, so a loss on both sides would go unnoticed. The independent CREATE VIEW / UNION ALL content checks are restored alongside the full comparison.

Follow-up 2, 205c9b1..185f582: stubs and live test CONFIRMED-RESOLVED. One remaining FINDING, recorded rather than fixed in bde5075: a CANCELLED query, or one FAILED for a reason other than a missing view, also raises a cause-less OperationalError and is still reported as absence. The reviewer confirmed this mapping predates the branch. Master mapped every OperationalError to NoSuchTableError; this branch only removed API failures from it. Distinguishing further would need a GetQueryExecution call on the failure path. I measured one: a missing view reports ErrorCategory 1 (SYSTEM) and ErrorType 1502, which does not single it out, so keying on it would be a guess. The comment now states the limit instead of implying a precise test, and the PR body lists it. No has-cause-but-failed-query path was found, sync or async.

Live after the repairs: a missing view gives NoSuchTableError on rest and pandas, and test_get_view_definition_across_cursor_types passes 5 of 5.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final independent follow-up, 185f582..8e90677 (same Codex reviewer, read-only): CONFIRMED-RESOLVED. There is no behavior change. The only edits are the rename, the docstring, and the new limit comment in get_view_definition(), and the decision logic of _is_fallback_error is unchanged. git grep finds no stale reference to the old name in tracked code, tests or docs. The new comment and docstring match the code: a failed or cancelled query raises a cause-less OperationalError and is read as absence, while API-layer failures keep their cause and propagate. The ErrorCategory/ErrorType measurement was taken as given, not re-derived.

# Athena returns the definition one line per row and blank lines as
# empty values, which are part of the definition.
return "\n".join(row[0] or "" for row in rows)

@reflection.cache
def get_columns(self, connection: Connection, table_name: str, schema: str | None = None, **kw):
Expand Down
Loading
Loading