From 4f2e5d9f577ec78c9bf6329fe47c10d14a80e14a Mon Sep 17 00:00:00 2001 From: Imran Siddique Date: Mon, 10 Aug 2026 08:46:28 -0700 Subject: [PATCH] docs(policy): hot-reload is implemented and inert in production The plan item said "design live policy reload, the interval is 0". The interval being 0 is not the whole story: reload is already built in PolicyStore.reload_if_stale, wired into PolicyEvaluator, and documented as a supported knob. It cannot swap a policy in any production configuration. Production requires CMCP_POLICY_HASH. That pinned hash is re-used on every reload, so the reload asks a question that answers itself: has the bundle changed, and does it still hash to the value it had before it changed. A genuinely changed bundle raises PolicyHashMismatch, gets swallowed by a bare except, and the old policy keeps being enforced with only a WARNING. The line that would install a new bundle is unreachable whenever a hash is pinned. Measured, not inferred: an allow-all to deny-all edit leaves the gateway enforcing allow-all. And because the failure path never advances _last_reload_at, 50 policy evaluations after the interval elapses produce 50 full bundle re-reads with hashing, on the enforcement hot path. So setting the documented knob in production buys a load amplifier and no policy update. No test caught this because every reload test constructs PolicyStore without expected_hash, which is the dev-mode configuration production never uses. Design doc only, no code. It states the real question a fix has to answer (if not a startup hash, what authorises a new bundle) plus what a reload does to TRACE evidence, and lays out five options with their costs. Also corrects STATUS.md, docs/configuration.md and the README, which all described the knob as merely disabled rather than as a trap. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 2 +- STATUS.md | 4 +- docs/configuration.md | 2 +- docs/spec/policy-hot-reload.md | 230 +++++++++++++++++++++++++++++++++ mkdocs.yml | 1 + 5 files changed, 235 insertions(+), 4 deletions(-) create mode 100644 docs/spec/policy-hot-reload.md diff --git a/README.md b/README.md index 8810768b..96c08810 100644 --- a/README.md +++ b/README.md @@ -178,7 +178,7 @@ catalog_path: catalog.json # approved tool catalog listen_addr: "127.0.0.1:8443" # tokenless dev mode is loopback-only; set CMCP_BEARER_TOKEN before binding wider max_response_size_bytes: 2097152 # 2 MB default -policy_reload_interval_seconds: 0 # 0 = disabled; restart required to update policy +policy_reload_interval_seconds: 0 # leave at 0; above 0 is a trap in production, see docs/spec/policy-hot-reload.md ``` Environment variables: diff --git a/STATUS.md b/STATUS.md index 2a205a2d..1547e462 100644 --- a/STATUS.md +++ b/STATUS.md @@ -12,7 +12,7 @@ picture is stated once. Developer Preview: interfaces may change before v1.0. | `attestation.enforcement_mode` | `enforcing` | | `attestation.staleness_policy` | `fail_closed` | | `attestation.validity_seconds` | `86400` | -| `policy_reload_interval_seconds` | `0` (disabled; policy change requires an enclave restart) | +| `policy_reload_interval_seconds` | `0` (disabled; policy change requires an enclave restart. Do not raise it: see the hot-reload row below) | | `attestation.allow_unmeasured_spawn` | `false` (a stdio server the catalog does not pin is not spawned) | | `attestation.required_provenance_kind` | `null` (server provenance is recorded, not enforced) | @@ -34,7 +34,7 @@ picture is stated once. Developer Preview: interfaces may change before v1.0. | Transparency-log anchoring for TRACE Claims | v0.2 | Write and lookup. | | Server-side (provider) attestation | Not yet (Phase 2) | Phase 1 attests the gateway boundary only. | | Server provenance checking | Shipped | Consumes [server-provenance-v1](https://github.com/agentrust-io/trace-spec/blob/main/spec/server-provenance-v1.md) records: verifies the signature against a configured publisher key, then compares the record's tool-catalog hash against the tools the server advertises **to this gateway**. Five outcomes reach the audit chain and none of them is silent: `verified`, `catalog-mismatch` (the document is fine and the server is not), `invalid`, `unchecked` (verified but the tool list was unavailable, so the comparison that matters never ran), `absent`. Absence is recorded and non-fatal by default, because almost no MCP server has a record and a gateway that refuses to route without one gets disabled on first contact. Set `attestation.required_provenance_kind` for a floor. | -| Real-time policy update without enclave restart | Not yet | `policy_reload_interval_seconds` is `0`; a policy change requires a restart. | +| Real-time policy update without enclave restart | Not yet, and the knob that looks like it is a trap | `policy_reload_interval_seconds` defaults to `0`, and a policy change requires a restart. Setting it above `0` does **not** buy hot-reload in production: the reload re-validates the new bundle against the hash pinned at startup, so any genuinely changed bundle fails and the old policy keeps being enforced. Worse, the failure path does not advance the interval, so every subsequent tool call re-reads and re-hashes the whole bundle on the enforcement path. Leave it at `0`. See [policy-hot-reload.md](docs/spec/policy-hot-reload.md) for the measurements and the options. | | AARM R4 five decision types | Shipped, with caveats | ALLOW, DENY, MODIFY, STEP_UP, DEFER are recorded in the audit chain. MODIFY is recorded as `redact`, DEFER is classified but not asynchronously enforced, and the TRACE Claim still carries the pre-AARM vocabulary. See [LIMITATIONS.md](LIMITATIONS.md). | | AARM R8 telemetry export | Shipped | OpenTelemetry spans mirroring audit entries. Opt in with `CMCP_OTEL_ENABLED=1` and `pip install cmcp-runtime[otel]`; a no-op otherwise. Exports digests, never payloads. The audit chain stays authoritative. | | AARM R2/R3 declared intent | Not implemented | cMCP takes no declared-intent input, so the intent-alignment half of R2 and R3 is unmet. Adding one changes the MCP-facing surface. | diff --git a/docs/configuration.md b/docs/configuration.md index a739003f..1fb65bef 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -105,7 +105,7 @@ All fields are optional as a group. If `path` is set, `trust_anchor_path` must a | `catalog_path` | string | `catalog.json` | Path to the JSON tool catalog. Path traversal (`..` components) is rejected. | | `listen_addr` | string | `0.0.0.0:8443` | Address and port the gateway binds to. Default is `127.0.0.1:8443` in tokenless `CMCP_DEV_MODE=1`, otherwise `0.0.0.0:8443`. Tokenless dev mode requires loopback (e.g., `127.0.0.1:8443`, `localhost:8443`, `[::1]:8443`). Wildcard, LAN, public, and non-loopback hostname binds require `CMCP_BEARER_TOKEN`. | | `max_response_size_bytes` | integer | `2097152` | Maximum tool response size in bytes (2MB). Must be a positive integer. Responses exceeding this limit are rejected before inspection. | -| `policy_reload_interval_seconds` | integer | `0` | Interval in seconds between automatic Cedar bundle reloads. `0` disables automatic reload. See note in the full example above. | +| `policy_reload_interval_seconds` | integer | `0` | Interval in seconds between automatic Cedar bundle reloads. `0` disables automatic reload, and **`0` is the only value to use today.** Above `0` the reload re-validates the new bundle against the hash pinned at startup (`CMCP_POLICY_HASH`), so a bundle that actually changed is rejected and the old policy stays in force; the failed reload also does not advance the interval, making every later tool call re-read and re-hash the bundle on the enforcement path. It works only under `CMCP_DEV_MODE=1`, where no hash is pinned. See [Policy Hot-Reload](spec/policy-hot-reload.md). | ## Environment variables diff --git a/docs/spec/policy-hot-reload.md b/docs/spec/policy-hot-reload.md new file mode 100644 index 00000000..b0cf2bdc --- /dev/null +++ b/docs/spec/policy-hot-reload.md @@ -0,0 +1,230 @@ +# Policy Hot-Reload + +**Document status:** Design proposal, nothing built +**Applies to:** cMCP Runtime gateway (`PolicyStore`, `startup`) +**Related config:** `policy_reload_interval_seconds` + +--- + +## Summary + +Hot-reload is not missing. It is implemented in `PolicyStore.reload_if_stale`, wired +into `PolicyEvaluator`, and documented as a supported knob — and **it cannot swap a +policy in any production configuration.** This document records why, measures what +the current code does instead, and lays out the options for fixing it. No code +changes accompany it. + +`STATUS.md` says real-time policy update is "Not yet" because +`policy_reload_interval_seconds` is `0`. That reads as "unimplemented, default off". +The truth is worse and more specific: it is implemented, it is off by default, and +turning it on in production buys a warning log line every request instead of a +policy update. + +## What is actually there + +`PolicyStore` (in `policy/bundle.py`) holds the active bundle behind an `RLock`. +`PolicyEvaluator.evaluate` calls `reload_if_stale()` on the way in. Once the +interval has elapsed, the store re-reads the bundle from disk and swaps it in if +the hash changed, keeping the current bundle if anything raises. + +That is a reasonable poll-and-swap design. The problem is the trust anchor it +re-uses. + +## The defect: the pin and the reload contradict each other + +At startup the gateway **requires** `CMCP_POLICY_HASH` unless `CMCP_DEV_MODE=1` +(POLICY-001/#137: without a pinned hash, a tampered bundle loads silently). That +pinned hash is handed to `PolicyStore` as `expected_hash` and then re-used on +**every reload**: + +```python +new_bundle = load_policy_bundle(self._bundle_path, self._expected_hash) +``` + +`load_policy_bundle` raises `PolicyHashMismatch` when the bundle on disk does not +match the hash it was given. So the reload path asks a question that answers +itself: *"has the bundle changed, and does it still hash to the value it had before +it changed?"* A bundle that has genuinely changed always fails. A bundle that has +not changed always passes and swaps nothing. + +The line that would install a new bundle: + +```python +if new_bundle.bundle_hash != self._bundle.bundle_hash: +``` + +is unreachable whenever `expected_hash` is set, because the call above it has +already raised. + +**Measured, not inferred.** With `expected_hash` set to the startup hash, an +operator edit from allow-all to deny-all produces: + +``` +startup bundle hash : sha256:fdb2e39e... +new on-disk hash : sha256:e1426d71... +reload_if_stale() : False +active bundle hash : sha256:fdb2e39e... <-- unchanged +``` + +The gateway keeps enforcing allow-all. The operator's deny-all edit does not take +effect, and the only signal is a `WARNING`. + +### Why no test caught it + +Every reload test in `tests/unit/test_policy_bundle.py` constructs `PolicyStore` +**without** `expected_hash`: + +```python +store = PolicyStore(bundle=old_bundle, bundle_path=str(bundle_dir), + reload_interval_seconds=1) +``` + +That is the dev-mode configuration, where reload does work. +`test_policy_store_bundle_swap_on_hash_change` passes and proves the swap logic is +correct — in the one configuration production never uses. The pinned-hash case, +which is the only case a deployment runs, is untested. + +### Second defect: it re-reads the bundle on every request, forever + +`reload_if_stale` advances `_last_reload_at` only on the success path. The failure +path returns without touching it, so the staleness check stays true and the next +call tries again immediately. Since the reload in production *always* fails, the +interval stops being an interval. + +Measured: 50 policy evaluations after the first interval elapsed produce **50 full +bundle re-reads from disk**, each one reading every policy file and computing a +SHA-256 over the canonical bundle, on the request hot path. + +So the production effect of setting `policy_reload_interval_seconds: 60` is not +"policy updates every minute". It is "after one minute, every tool call performs +full bundle I/O and hashing, and the policy never changes". That is a +self-inflicted load amplifier on the enforcement path, reached by following the +documented configuration. + +### Third, smaller: the error text is from the wrong lifecycle + +The swallowed exception logs `Policy bundle hash mismatch: gateway will not start` +during steady-state operation. The gateway is already running and is not going to +stop. `except Exception` also catches genuine faults (unreadable file, malformed +Cedar, permissions) and files them all under the same warning, so the log cannot +distinguish "operator edited the policy and the pin now disagrees" from "the disk +is failing". + +## The real design question + +Hot-reload and hash pinning are not accidentally in tension — they want different +things: + +- **Pinning a hash** says *the policy is exactly this artifact, decided before the + process started, and nothing may change it afterwards.* That is what makes a + policy bundle attestable: the hash goes into the TRACE claim, and a verifier can + check that the gateway enforced the bundle it said it did. +- **Hot-reload** says *the policy may change while the process runs.* + +A design that wants both must answer: **when a new bundle arrives, what authorises +it?** A hash fixed at startup cannot, by construction. Something else has to. + +There is a second question that follows immediately and matters just as much for +cMCP specifically: **what does a reload do to evidence?** A TRACE claim names the +policy bundle hash the call was evaluated under. If the bundle can change +mid-process, then claims from one process carry different bundle hashes, and every +consumer of those claims needs to cope with that. Any option below has to say what +the claim records and how a verifier reconstructs which policy was live for a +given call. + +## Options + +### A. Pin a signing key, not a bundle hash + +The bundle carries a signature over its own manifest; the gateway pins the +**public key** allowed to sign policy. A reload verifies the new bundle's +signature rather than its hash. + +- Trust moves from "this exact artifact" to "any artifact this authority + approves", which is what actually makes runtime change safe. +- Fits the existing manifest (`author_identity`, `commit_sha` are already there, + unsigned) and the direction the rest of the stack has taken. +- Cost: key management, revocation, and a rollback story — a validly signed *older* + bundle is a downgrade attack unless the manifest carries a version that must + increase monotonically. +- Evidence: the claim records the bundle hash *and* the signer plus the bundle + version, so a verifier can check both what ran and who authorised it. + +### B. Pin a set of acceptable hashes + +`CMCP_POLICY_HASH` becomes a list. A reload accepts any bundle whose hash is in the +allowlist. + +- Smallest change from what exists, and keeps the "exactly these artifacts" model. +- No new cryptography and nothing to revoke. +- Cost: every policy change still needs the operator to restart to extend the + allowlist, so it does not deliver hot-reload — it only lets a fleet roll between + a known set of policies without restart. Useful for staged rollout and + fast rollback; not an answer to "we need to tighten a policy right now". + +### C. Re-read the pin from a trusted source at reload time + +The expected hash comes from somewhere the gateway can re-consult — a separate +hash file, a control plane, a transparency log — instead of a startup env var. + +- Genuine hot-reload, and the authority for a policy change stays outside the + gateway. +- Composes with the transparency direction: the pin could be an entry a verifier + can independently look up. +- Cost: introduces a runtime dependency on that source and its own trust + question. If the hash file sits next to the bundle and is writable by whoever + writes the bundle, it authorises nothing. + +### D. Operator-triggered reload that supplies the new hash + +No polling. An admin action (signal, authenticated endpoint) hands the gateway the +new expected hash and it reloads once. + +- Nothing changes under a running request without a human or a deploy pipeline + asking, which is the most predictable behaviour and the easiest to audit. +- Removes the hot-path `reload_if_stale()` call from `evaluate` entirely, and with + it the load amplifier. +- Cost: a new authenticated control surface on the gateway, which is attack + surface on the enforcement component. Less convenient than polling. + +### E. Keep it development-only, and make that honest + +Accept that a pinned, attestable policy and runtime mutation do not belong in the +same deployment. Reload is permitted only when `CMCP_DEV_MODE=1`; configuring +`policy_reload_interval_seconds > 0` together with a pinned hash is a **startup +config error** rather than a warning at request time. + +- Honest, costs nothing to build, and removes the failure mode entirely. +- Matches how the rest of the runtime treats this class of thing: fail at + construction on a configuration that could never work. +- Cost: cMCP keeps telling operators that a policy change needs an enclave + restart, which for a confidential gateway means an attestation cycle. That is a + real operational burden and the reason this item is on the list at all. + +## Recommendation is not the point yet, but two things are not optional + +Whichever direction is chosen, two fixes stand on their own and do not depend on +it: + +1. **A configuration that cannot work must not start.** `expected_hash` set + together with `reload_interval_seconds > 0` is, today, guaranteed-inert. Under + option E that pairing is the config error; under A/B/C/D it stops existing. + Either way it must never again be a thing an operator can switch on and believe. +2. **The interval must be honoured on failure too.** `_last_reload_at` has to + advance whether or not the reload succeeded, or a failing reload turns into + per-request bundle I/O on the enforcement path. This is worth fixing even if + hot-reload is removed, because the same shape will reappear in the next thing + that polls. + +And the test gap generalises past this feature: a code path whose only tests +construct it in the configuration production never uses is a path with no tests. +The reload tests should be parameterised over pinned and unpinned, so whichever +option lands is exercised the way it will actually be deployed. + +## Not in scope + +- Catalog hot-reload. `CMCP_CATALOG_HASH` has exactly the same pin, and + `load_catalog` the same shape, so whatever is decided here should be applied + there deliberately rather than by copy. It is not analysed in this document. +- What a reloaded policy means for an in-flight session that has already been + admitted under the previous bundle. diff --git a/mkdocs.yml b/mkdocs.yml index d77ae140..fd7dac80 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -182,6 +182,7 @@ nav: - Error Codes: spec/error-codes.md - Failure Modes: spec/failure-modes.md - Threat Model: spec/threat-model.md + - Policy Hot-Reload: spec/policy-hot-reload.md - Phase 2 Server: spec/phase2-server.md - Testing: - Benchmarks: testing/benchmarks.md