Skip to content

Investigate persistent metadata throttling during SQLAlchemy reflection #780

Description

@laughingman7743

Problem

SQLAlchemy reflection can still fail with Athena metadata throttling after the request-reuse and concurrency changes in #777.
Restoring reflection coverage in #771 exposed the problem during the compliance suite; the root cause and practical impact outside CI are not yet established.
Track the remaining investigation separately from the weekly database sweep and development-tooling work in #779.
This blocks completion of the prerequisite metadata fix for the ARRAY work in #755.

Current implementation and evidence

Continue from #777, currently Draft at ea4ce1fc04c922592ef9f6b0fad46702c894598c.
Do not discard the existing implementation or start a replacement PR without a concrete reason.

  • Positive table metadata is shared within one SQLAlchemy Inspector across listings and column/comment/table-option reflection. Catalog and schema identity, cache invalidation, and preservation of earlier results have regression coverage.
  • Fresh CREATE/DROP existence checks remain uncached. A profiled synchronous baseline attributed 378 of 643 GetTableMetadata SDK calls to these checks.
  • Wrapped throttling, access denial, and unknown metadata failures must propagate rather than establish or cache table absence.
  • CI SDK retries are bounded to standard mode with three attempts. SQLAlchemy matrices run at most two Python versions concurrently, with eight workers per job; this does not limit traffic from other PRs.
  • CI run 35448063098 completed on that head: all five regular PyAthena jobs passed; SQLAlchemy 3.10, 3.11, and 3.12 failed, while 3.13 and 3.14 passed. The dependent async suite was skipped.
  • Catalog cleanup improved database enumeration, but metadata throttling recurred afterward. The database backlog is not established as its cause.
  • Earlier full local synchronous and asynchronous baselines passed, but used a different AWS account from CI. They do not establish that the current head passes in the CI account. Use the current local .env configuration and verify the target before interpreting new results.
  • Independent review findings about error-propagation regression coverage and documentation were repaired in the current head. The final independent follow-up remains incomplete.

Questions to resolve

  1. Which API calls, callers, retry layers, and concurrency patterns cause the remaining failures in the intended test account? Separate Inspector reflection from fresh existence checks and distinguish client calls from actual SDK attempts.
  2. Does comparable reflection traffic fail in normal application use, or only under the compliance suite's load? Measure request counts, elapsed time, and failure rates under controlled concurrency before proposing further retries or a quota increase.
  3. Can additional request reuse reduce traffic without changing metadata freshness, cross-catalog behavior, or error visibility?
  4. Would a limited information_schema path help bulk column reflection? Current Athena checks found column comments, type parameters, and partition markers, but not all table comments, location, SerDe, or table properties used by the dialect. A full replacement would lose information, and information_schema itself can fail with large catalogs. Any experiment must compare semantics and traffic as well as latency.

Reproduction and validation

Use the existing #777 worktree/branch and repository workflow.
Run just worktree-env in a worktree, then use uv run --env-file .env ... to load the intended AWS configuration without printing credentials.
Do not overlap local live-Athena tests with CI or other benchmark/test workloads.

First recheck the current PR head, CI logs, and any active test runs.
Then reproduce with existing HasTableTest cases and temporary API instrumentation when useful:

just lint
uv run --env-file .env pytest -n 1 tests/sqlalchemy/test_suite.py::HasTableTest -q
uv run --env-file .env pytest -n 1 tests/sqlalchemy/test_suite.py::HasTableTest --dburi async -q

These targeted commands are starting points for investigation, not evidence that the entire current suite passes.
After a repair, run the required synchronous and asynchronous suites sequentially:

just lint
uv run --env-file .env just test sqla
uv run --env-file .env just test sqla-async

Use the existing test classes and fixtures; do not add a TestMetadataReflection class or a parallel mock-heavy framework.
Real S3 Tables absence handling has been checked; Lambda/Hive federated catalogs remain unverified.
Preserve that distinction when reporting compatibility.

Code and history pointers

The historical PR descriptions do not prove that earlier Athena versions lacked type parameters: the old client also stripped them.
Preserve this distinction when evaluating a new information_schema experiment.

Completion criteria

  • Document a reproducible failure scenario and the evidence supporting the chosen repair, including limits on production-impact claims.
  • Preserve cache invalidation, catalog/schema isolation, fresh existence checks, and missing-object versus failed-request semantics.
  • Do not mask unexplained failures by increasing retry budgets or weakening tests.
  • Pass applicable local validation and current aggregate CI, including the async suite, before marking Fix metadata reflection throttling and reuse listed metadata #777 Ready.
  • Record two distinct self-reviews and collect the independent follow-up required by the repository workflow.

Related work

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions