Skip to content

Latest commit

 

History

History
156 lines (118 loc) · 11.6 KB

File metadata and controls

156 lines (118 loc) · 11.6 KB

CLAUDE.md

Guidance for AI coding assistants working in this repository.

What this project is

aws_resource_validator is a Python package that ships two kinds of artefacts:

  1. Resource validator — regex-backed validators + sample-name generators for AWS resource names (ARNs, bucket names, etc.), keyed by service. Single committed file: aws_resource_validator/class_definitions.py.
  2. Pydantic v2 models — one sub-package per AWS service under aws_resource_validator/pydantic_models/<service>/ mirroring the TypedDict surface of mypy_boto3_<service>. Consumers write from aws_resource_validator.pydantic_models.s3.s3_classes import BucketTypeDef.

Both artefacts are generated by code under aws_resource_validator/generator/ but are committed so consumers don't need network access or stub packages at install time.

Python target: 3.11+. Pydantic: v2.8+.

Repository layout

aws_resource_validator/
├── core/                              # Runtime library (small, stable).
│   ├── api_object.py                  # Frozen APIObject dataclass.
│   ├── service.py                     # Container of APIObjects.
│   ├── registry.py                    # Mapping of service name -> Service.
│   ├── base_validator_model.py        # The ONE BaseValidatorModel + EventStream[T].
│   └── naming.py                      # to_snake_case, to_pascal_case, safe_identifier.
│
├── class_definitions.py               # GENERATED by Pipeline A. Do not edit.
│
├── pydantic_models/                   # GENERATED by Pipeline B. Do not edit.
│   └── <service>/
│       ├── __init__.py                # Re-exports all public classes.
│       ├── <service>_classes.py       # Pydantic BaseModels.
│       ├── <service>_constants.py     # Literal type aliases.
│       └── _manifest.json             # Used by tests/generated/.
│
└── generator/                         # Build-time only; never imported at runtime.
    ├── cli.py                         # `arv-generate` Typer CLI.
    ├── common/                        # logging, atomic I/O, black wrapper, Jinja env cache.
    ├── pipeline_a/                    # GitHub fetch -> APIRegistry -> class_definitions.py.
    └── pipeline_b/                    # mypy_boto3 stubs -> ServiceModelIR -> per-service files.

tests/
├── unit/                              # Hand-written; grouped by source module.
├── integration/                       # Cross-module smoke tests.
└── generated/                         # ONE file, parametrized from manifests (~65k cases).

Commands

# Install (Poetry).
poetry install

# Unit + integration tests (fast).
pytest tests/unit tests/integration

# Full parametrized suite over every generated Pydantic class (~65k tests, parallelized).
pytest tests/generated -n auto

# Lint & type-check.
ruff check aws_resource_validator tests
mypy aws_resource_validator/core aws_resource_validator/generator

# Regenerate everything.
arv-generate all                      # Requires GITHUB_TOKEN + mypy_boto3 stubs in venv.
arv-generate pipeline-a               # Just class_definitions.py.
arv-generate pipeline-b               # Just pydantic_models/ tree.
arv-generate pipeline-b --only acm    # Single service (fast iteration).

arv-generate pipeline-a makes authenticated GitHub API calls and only needs to run when AWS adds a service or updates a shape pattern. arv-generate pipeline-b is purely local — it walks mypy_boto3_* packages in the venv's site-packages and should be re-run whenever those stubs are upgraded.

Critical invariants

  1. Generated trees are committed but never hand-edited. class_definitions.py and everything under pydantic_models/<service>/ comes out of the CLI. Editing them is a recipe for confused diffs. These paths are excluded from ruff, mypy, and coverage.
  2. core/ has no generator imports. The generator depends on core; core never depends on the generator. Violations break the runtime library's pip install.
  3. Pipeline B is single-pass. Every extractor (literal_extractor, typeddict_extractor, service_annotator) produces IR; the emitter renders templates from IR. Do NOT add a post-processing pass that rewrites emitted files — the previous codebase had one (annotate_boto3_classes.py + replace.zsh) and it was a source of silent bugs.
  4. Generated *_classes.py must NOT use from __future__ import annotations. Pydantic v2 with __future__ annotations needs model_rebuild() on every class, which we don't emit. The generator output is Optional[X]-style typed and must be evaluable at class-body time.
  5. Source order matters in generated files. The upstream mypy_boto3_* stubs are topologically sorted; ServiceModelIR.items preserves that order in a single tuple so forward references resolve at import time. If you split classes from aliases into two tuples, you'll get NameError on import.

Gotchas the generator handles (don't re-break these)

  • PEP 585 lowercase builtins (list[X], dict[K, V], set[X], tuple[X, ...], type[X]) are rewritten to their typing counterparts (List[X], Dict[K, V], ...) by TypeResolver._TYPING_ALIASES. This prevents field-shadowing crashes — e.g. iotsitewise has a TypedDict with a field literally named list annotated list[dict[str, Any]]; with the builtin, the field name shadows the constructor at class-body eval time.
  • lambda is an AWS service AND a Python keyword. safe_identifier() suffixes it — the emitted package is aws_resource_validator.pydantic_models.lambda_.lambda__classes.
  • Field names may be Python keywords (or, and, from) or contain hyphens (detail-type). typeddict_extractor._safe_field_name produces a Python-safe identifier and the emitter adds Field(alias=wire_name) so the JSON representation is unchanged.
  • EventStream[T] classes are not Pydantic models. They inherit from botocore.eventstream.EventStream (via a generic wrapper in core/base_validator_model.py) and intentionally lack model_validate / model_json_schema. The generated test suite skips them after the issubclass check.
  • Pydantic extra="allow" on BaseValidatorModel means unknown keys are silently preserved. This is intentional forward-compat with future boto3 releases; strict validation must walk the declared field list.
  • Callable fields break model_json_schema(). S3 has a few (e.g. BucketDownloadFileRequestTypeDef.Callback). The parametrized test suite catches PydanticInvalidForJsonSchema as an expected outcome — don't add a generic except that hides real failures.

Testing philosophy

  • Every hand-written public function has at least one unit test.
  • Every generated Pydantic class has a parametrized test through tests/generated/test_pydantic_models.py, driven by _manifest.json files emitted alongside each <service>_classes.py. The contract checked per class:
    1. Subclass of BaseValidatorModel or EventStream.
    2. model_json_schema() either succeeds or raises PydanticInvalidForJsonSchema (the one known-acceptable failure).
    3. model_validate({}) either succeeds (all fields optional) or raises ValidationError (required fields present) — never anything else.
  • Run the full suite with pytest-xdist -n auto (~25s for 65k tests on modern hardware).

Regenerating after mypy_boto3_* stubs change

# Update stubs in the venv.
poetry update boto3-stubs

# Wipe existing output and regenerate.
find aws_resource_validator/pydantic_models -mindepth 1 -maxdepth 1 -type d -exec rm -rf {} +
arv-generate pipeline-b

# Verify.
ruff check aws_resource_validator/core aws_resource_validator/generator tests
pytest tests/unit tests/integration
pytest tests/generated -n auto

If regeneration produces new failures, the fix almost always goes in one of:

  • generator/pipeline_b/type_resolver.py (new AST shape encountered).
  • generator/pipeline_b/typeddict_extractor.py (new field-name quirk).
  • generator/pipeline_b/templates/classes.py.j2 (new symbol needs importing).

Regeneration is deterministic — arv-generate pipeline-b && git diff --exit-code should be clean on a fresh stub set.

Legacy shims currently in the tree

  • aws_resource_validator/models.py — re-exports APIObject, Service, APIRegistry from core/. Required until class_definitions.py is regenerated via Pipeline A (which needs a GITHUB_TOKEN).
  • core/registry.py::add_service and core/service.py::add_api_object(*args) — accept the legacy two-arg calling convention used by the currently-committed class_definitions.py. Both are scheduled for removal once Pipeline A regenerates that file.

When Pipeline A is re-run, delete models.py, add_service, and the variadic form of add_api_object together.

Release packaging

The runtime package is split across a family of wheels:

  • aws-resource-validator — core + generator code. Ships no pydantic_models/ content (excluded via tool.poetry.exclude).
  • aws-resource-validator-<svc> — one per popular AWS service listed in scripts/release/popular_services.txt. Ships only the pydantic_models/<svc>/ subtree and depends on the main wheel. There is also one wheel per service directory (423 in total), so every extras key resolves.
  • aws-resource-validator-<shard> — metapackage per domain shard from scripts/release/shards.toml. Empty wheel whose Requires-Dist pulls in every service wheel in that shard.
  • Optional deps + extras for the main wheel are generated by scripts/release/sync_extras.py between # --- BEGIN GENERATED ... --- markers in the root pyproject.toml. Never hand-edit inside those markers; run python -m scripts.release.sync_extras --write. CI enforces synchronization via --check.
  • All per-service and shard wheels use hatchling so pydantic_models/ can stay a PEP 420 namespace package. The single blocker for the namespace is any __init__.py under pydantic_models/ — don't reintroduce one, and don't let any wheel ship aws_resource_validator/__init__.py except the main wheel. scripts/release/verify_wheels.py enforces this.

Conventions

  • Pytest for tests (not unittest.TestCase).
  • Ruff for lint + format (config in pyproject.toml). No .flake8 or .pylintrc.
  • Mypy strict on core/ and generator/; the generated trees are excluded.
  • Frozen dataclasses + slots for value types; Mapping / MappingProxyType for read-only views.
  • Logging via generator/common/logging.get_logger(__name__) — never print().
  • File writes via generator/common/io.write_text_atomic — crash-safe mkstemp + rename.
  • Jinja2 environments from generator/common/templates.build_environment(path) — cached per directory.

What NOT to do

  • Don't commit anything under pydantic_models/<service>/ or class_definitions.py by hand — regenerate.
  • Don't add runtime imports from aws_resource_validator.generator.* in core/ or in generated files.
  • Don't bypass write_text_atomic for generator output; partial files caused the old replace.zsh bug chain.
  • Don't reintroduce from __future__ import annotations in the classes.py.j2 template.
  • Don't introduce dynamic code execution in the generator or runtime (exec, eval, compile, unsafe deserialisation formats, subprocess with shell=True). All code generation goes through Jinja2; all code reading goes through ast.parse.
  • Don't add an __init__.py under aws_resource_validator/pydantic_models/ — it's a PEP 420 namespace package and any regular-package marker breaks cross-wheel imports from the aws-resource-validator-<svc> family.