Rehearse every critical change. Trust only proven agents. Execute with evidence.
ChangeMesh is a policy-governed agent fleet that safely rehearses and executes high-risk enterprise architecture changes. It treats an enterprise change as a long-lived distributed transaction, not a chat session. Using the Google Agent Development Kit and Gemini, it discovers dependencies, rehearses migrations in a shadow environment, and compresses weeks of manual coordination into a single, tamper-evident Change Evidence Passport.
- Final Demo Video (3:24): https://www.youtube.com/watch?v=l5JYjOTCw1E
- Live Demo (Cloud Run): https://changemesh-p24-e2e-764732742797.europe-west3.run.app
- Judge Quick-Start Guide:
docs/JUDGE_START_HERE.md - Critical Evidence Boundary: The hosted current workflow runs in SIMULATION mode. The real GitHub Draft PR shown in the demo is historical LIVE_WRITE evidence (PR #2), was intentionally never merged, and is not presented as the output of the current simulation.
| Live Cloud Dashboard (Preferred Cover) | ShadowLab Rehearsal Twin |
|---|---|
![]() |
![]() |
| Approval Compression Decision Card | Evidence Passport & Cloud Proofs |
![]() |
![]() |
For complete captions, metadata, and cryptographic SHA-256 digests of all six presentation screenshots, see
docs/P-31.06_SCREENSHOT_MANIFEST.md.
Important
Frozen hackathon submission candidate; release-blocking fixes only.
Core implementation phases P-00 through P-31 (including P-31.06 final screenshot pack) and P-32.01 (submission freeze) are complete (DONE). 1,900+ automated tests pass with zero failures. Google Cloud Run service deployed (changemesh-p24-e2e-00004-djv in europe-west3).
- Status: Submission freeze active; final validation and candidate release in progress (P-32).
- Evidence Boundary: The hosted current workflow runs in SIMULATION mode. The real GitHub Draft PR shown in the demo is historical LIVE_WRITE evidence (PR #2), was intentionally never merged, and is not presented as the output of the current simulation.
- Hackathon: All Things Agentic Hackathon
- Primary category: Fortified Enterprise Fleet (See
docs/CATEGORY_MAPPING.mdfor concrete architectural mapping) - Required model path: Gemini 3.5 or newer through Vertex AI or the Gemini API
- Primary agent framework: Google Agent Development Kit (ADK)
- Required cloud proof: Google Cloud deployment and runtime evidence
- Planned core services: Agent Runtime/Platform + Cloud Run for supporting services, Firestore (Operational saga state), Pub/Sub
- Target enterprise services: Agent Runtime/Platform (AVAILABLE), Agent Platform Memory Bank + ChangeMesh Memory Trust Layer (DEFERRED), Agent Registry (AVAILABLE), Agent Identity (SPIFFE-based) + ChangeMesh Capability Passport (PERMISSION_BLOCKED), Agent Gateway (networkservices) + ChangeMesh Policy Guardian (AVAILABLE), Model Armor (PERMISSION_BLOCKED), ADK OpenTelemetry -> Cloud Logging/Trace (AVAILABLE)
A change that looks small in one system can cross source code, database schemas, data pipelines, dashboards, APIs, security policies, ownership boundaries, and release processes.
For example:
Rename
customer_idtoaccount_idacross the billing platform.
A conventional coding agent can change files. It usually cannot prove that it:
- found every downstream dependency;
- preserved backward compatibility;
- recovered from partial failure;
- used only authorized tools and data;
- resumed safely after days or weeks;
- distinguished real evidence from model confidence;
- escalated only the irreducible human decision;
- left a trustworthy handoff for the next agent or team.
ChangeMesh treats an enterprise change as a long-lived distributed transaction, not a chat session. (See Product Brief for full buyer, operator, and wedge definitions).
Every change moves through eight explicit stages:
- Discover — find relevant agents, tools, repositories, owners, and dependencies.
- Qualify — verify the exact agent revision through a Capability Passport.
- Rehearse — run policy-defined failure, attack, recovery, and stale-context scenarios in ShadowLab.
- Ground — load only trusted, scoped, non-expired institutional memory.
- Authorize — assign the smallest safe autonomy envelope through the Reversibility Gate.
- Execute — perform asynchronous work through idempotent, recoverable steps.
- Prove — collect deterministic evidence, traces, approvals, blocked actions, and
NOT_RUNstates. - Certify — seal the result in a Change Evidence Passport.
ChangeMesh is designed as human-on-the-loop, not approval-heavy human-in-the-loop software.
The fleet should autonomously perform reversible and policy-approved work such as dependency analysis, planning, branch creation, migration and rollback artifact generation, tests, rehearsal, retry, draft PR creation, evidence collection, and handoff preparation.
Human attention is reserved for actions where organizational authority is required, such as irreversible production mutation, sensitive-data movement, privilege expansion, protected-branch merge, or production deployment.
Instead of many meetings, messages, and repeated explanations, ChangeMesh produces one bounded decision card containing:
- what has already been completed;
- which evidence passed or failed;
- what remains uncertain;
- the smallest requested authority;
- the recommended safe decision;
- the effect of approval or rejection.
(See Outcome Contract for strict metrics on how Approval Compression reduces human touches).
Before a critical real action, ChangeMesh rehearses the workflow against controlled tools and synthetic enterprise context.
Initial scenarios:
- normal migration;
- GitHub/API
503recovery; - partial migration interruption;
- stale approval detection;
- indirect prompt injection;
- missing rollback proof;
- agent restart and resume;
- downstream client compatibility failure.
A failed required rehearsal denies real execution authorization until the workflow is corrected and re-evaluated.
Agent discovery is not treated as trust. Each exact agent revision receives a passport recording declared capabilities, scenario results, authorized data/action classes, evidence hashes, validity, and revocation.
The orchestrator must route a task to a proven revision, not merely to an agent that claims a matching skill.
Long-term memory is typed and governed. Initial memory classes:
OBSERVED_FACTHUMAN_DECISIONAGENT_INFERENCEASSUMPTIONPOLICYFAILED_APPROACHUNVERIFIED_EXTERNAL_INPUT
Decision-relevant memory must carry source, scope, timestamp, validity, sensitivity, and evidence linkage. Stale, contradictory, untrusted, or injection-suspected memories are quarantined rather than silently reused.
Autonomy is determined by impact and reversibility, not by one global permission switch.
Initial decision classes:
AUTO_EXECUTEAUTO_EXECUTE_AND_NOTIFYREHEARSE_THEN_EXECUTEHUMAN_AUTHORITY_REQUIREDBLOCKED
The final passport binds mission, agent/tool identities, delegation and event chain, trusted context hashes, changed-file manifest, tests, rehearsals, blocked and NOT_RUN actions, authority decisions, semantic evaluation, and final integrity hash.
ChangeMesh deploys six Google ADK agents total: the ChangeMesh Orchestrator plus five specialized worker agents.
| Agent | Primary responsibility | Default authority |
|---|---|---|
| ChangeMesh Orchestrator | Goal interpretation, dynamic routing, 8-stage saga coordination, recovery | Coordinate; no unrestricted production mutation; durable workflow state owned by Firestore Saga |
| Impact Scout | Repository, dependency, lineage, ownership, and conflict analysis | Read-only |
| Policy Guardian | Privacy, prompt injection, identity, tool and data policy evaluation | Block or constrain; no implementation writes |
| Migration Engineer | Safe expand–migrate–contract artifacts and tests | Write only to scoped branch/worktree |
| Evidence Auditor | Independent mission–change–test alignment review | Read-only; cannot rewrite deterministic facts |
| Release Steward | Draft PR, decision packet, passport, and handoff | Reversible release preparation only |
The frozen canonical end-to-end scenario is a synthetic Acme Billing schema migration:
billing_accounts ADD COLUMN payment_tier VARCHAR(32)without breaking downstream clients.
Expected demonstration:
- User provides one goal (
billing_accounts ADD COLUMN payment_tier VARCHAR(32)). - ChangeMesh discovers and qualifies required agent revisions.
- Agents work asynchronously across repository and metadata context.
- Direct destructive mutations are blocked; non-destructive expand strategy is synthesized.
- ShadowLab validates compatibility and rollback safety in synthetic twin environment.
- The fleet coordinates across Scout, Policy, Migration, Evidence, and Steward agents.
- Tests, migration artifacts, and rollback proof are synthesized in simulation; historical Draft PR #2 is displayed and replayed (
LIVE_WRITE / HISTORICALprovenance; zero fresh GitHub mutations during rehearsal). - A new session resumes from trusted memory.
- One compressed authority decision is requested only for the irreducible human boundary.
- A Change Evidence Passport is sealed; historical P-24 Google Cloud Trace (
RECORDED_CLOUD / HISTORICAL, tracec137e280da7d4f25ae08138649e6d374) is displayed (CURRENT deployed revision Cloud Trace isNOT_RUN).
The component dependency architecture is documented in docs/ARCHITECTURE.md. It defines:
- Component boundaries and canonical ownership
- Explicit dependency directions (inward dependency principle)
- Canonical planned package map
- Provider-neutral domain boundary
- Adapter replaceability contract
Important
Current Architecture & Implementation Status: Phases P-00 through P-31 (including P-31.06 final screenshot pack) and P-32.01 (submission freeze) are complete (DONE). Submission freeze is active; final validation and candidate release checks (P-32) are in progress. Current verified Google Cloud Run revision is changemesh-p24-e2e-00004-djv in europe-west3. Optional enterprise managed services without direct access remain honestly labeled NOT_RUN or AVAILABLE/NOT_RUN.
The competition runtime must not be presented as an Antigravity desktop automation.
- Antigravity: development environment and governed coding assistant.
- Google ADK: product multi-agent runtime and orchestration.
- Gemini 3.5+ via Vertex AI/Gemini API: runtime reasoning.
- Google Cloud: actual deployed backend and evidence source.
Where a target enterprise service is unavailable because of preview access, account, quota, or region:
- record the real limitation;
- keep the state
NOT_RUN; - use a clearly labeled local deterministic adapter only for development;
- never present the adapter as proof of the unavailable managed service.
Before changing code, every development agent must read:
AGENTS.mdCHANGEMESH_RULES.mdplans/CHANGEMESH_MASTER_EXECUTION_PLAN.mdAGENT_MEMORY_AND_LESSONS.mdAGENT_ARCHITECTURE_AND_PATTERNS.mdAGENT_ENVIRONMENT_AND_API.mdAGENT_USER_PREFERENCES.mddocs/HANDOFF.md
The project charter is already agreed. No Phase-0 interview is required. Questions are permitted only when a genuine blocking product decision cannot be derived from the frozen charter, repository evidence, policies, or memory.
ChangeMesh enforces a strict boundary between execution modes and result states (see docs/MODE_CONTRACT.md):
- Execution Modes:
FIXTURE,SIMULATION,RECORDED_CLOUD,LIVE_WRITE. Adapters execute the explicitly selected mode or fail; there is no silent fallback. Mode labels must be visible. - Evidence States: The result of the executed operation.
Simulation and fixtures are not live proof. Recorded-cloud is a replay of an actual past execution, not a live call. Live-write performs bounded real mutation (e.g., in a demo repository). Live-write does not automatically mean human approval is required, as organizational policy determines autonomy. The local in-memory event bus (LocalEventBus) carries explicit transport="LOCAL" and maps strictly to SIMULATION or FIXTURE mode; it cannot produce LIVE_WRITE or RECORDED_CLOUD evidence, preventing local simulation from being mistaken for Google Pub/Sub proof.
| State | Meaning |
|---|---|
PASS |
A named check or action actually completed successfully |
WARN |
Evidence exists but requires attention |
FAIL |
A named executed check failed |
NOT_RUN |
The check or integration was not executed |
SIMULATED |
The result came from an explicitly labeled simulation |
BLOCKED |
Policy prevented execution; the action remains NOT_RUN |
QUARANTINED |
Context or memory is excluded from decisions pending review |
Model opinions can evaluate semantic sufficiency. They cannot rewrite locked execution facts.
The binding, living roadmap is:
No implementation task is complete until the plan, architecture, memory, environment notes, README, handoff, and affected judge-facing documents are synchronized.
Clean-checkout reproducibility from a separate directory outside the canonical workspace has been verified under P-06.05 (docs/P-06.05_CLEAN_CHECKOUT_LOG.md).
- Python:
3.13.5(managed viauvor system CPython 3.13.5, pinned in.python-version) - uv:
0.11.28(pinned inpyproject.toml[tool.uv] required-version) - Git
Environment Tested: Windows 11 x86_64, PowerShell 7, CPython 3.13.5, uv 0.11.28, Git 2.52.0.
-
Clone the repository:
git clone https://github.com/zyganali-glitch/ChangeMesh.git cd ChangeMesh -
Synchronize dependencies (deterministic frozen install):
uv sync --frozen
-
Verify dependency consistency:
uv pip check
-
Run unit tests:
uv run python scripts/cmd.py unit
(Executes unit/contract tests across all implemented domain, agent, event, memory, capability, shadowlab, gate, policy, migration, audit, release, and saga modules with exit code 0; one ADK deprecation warning is recorded.)
- No
.envrequired: Local unit tests, schema validations, and command checks do not require.envor cloud credentials. - Safe template:
.env.exampleprovides the canonical environment structure with zero secret defaults. - Google Cloud Auth: Application Default Credentials (
gcloud auth application-default login) are required only when running explicitly authorized live Google Cloud operations in later phases. - Service-Account Keys: Service-account JSON key files are prohibited and strictly ignored by
.gitignore.
| Command | Action | Check Semantics | Baseline Result |
|---|---|---|---|
uv run python scripts/cmd.py unit |
Run unit tests | Local deterministic test suite (excluding live mutation tests) | PASS (deterministic suite green, 1 warning) |
uv run python scripts/cmd.py format |
Format check | Non-mutating (ruff format --check .) |
PASS (check-only) |
uv run python scripts/cmd.py lint |
Lint check | Non-mutating (ruff check ., zero --fix) |
PASS (check-only, 0 violations) |
uv run python scripts/cmd.py type-check |
Type-check | Non-mutating (mypy domain src integrations tests service_app.py) |
PASS (check-only, 0 issues) |
uv run python scripts/cmd.py integration |
Integration tests | Fails closed by default; requires explicit --live-write-danger |
FAIL_CLOSED (guarded real GCP mutations) |
uv run python scripts/cmd.py demo |
Synthetic E2E Demo | Runs local synthetic Acme Billing E2E (SIMULATION) |
PASS (deterministic local simulation) |
uv run python scripts/cmd.py e2e |
E2E command | Intentionally disabled in CLI; zero live mutation | NOT_RUN (verified via hosted /run-e2e & tests) |
uv run python scripts/cmd.py deploy |
Deploy command | Safety-disabled in canonical CLI; zero mutation | NOT_RUN (verified deployment provenance in P-28.03) |
uv run python scripts/cmd.py teardown |
Teardown command | Safety-disabled in canonical CLI to prevent accidental deletion | NOT_RUN (zero cloud mutation) |
ChangeMesh is engineered with defense-in-depth safety, deterministic policy guardians, and cryptographic audit ledgers. In adherence to transparent engineering principles:
- Non-Certification Notice: ChangeMesh provides governance readiness and automated safety verification; it is not an officially certified substitute for accredited third-party compliance audits (e.g., SOC 2, HIPAA, PCI-DSS, FedRAMP).
- Deterministic Code Primacy: Generative AI models (Gemini 3.6 Flash) provide semantic advice and structural artifact generation. Deterministic code gates and policy guardians maintain absolute veto power over model opinions.
- Draft-Only Pull Requests: Autonomous agents are restricted to draft pull requests; direct branch push and automated pull request merging are structurally forbidden.
- Cloud Isolation: High-risk fault rehearsals and scenario attacks are strictly isolated within ShadowLab simulated environments.
(For detailed threat analysis, residual risks, and mitigation controls, see Threat Model and Non-Goals).
The initial product wedge is high-risk schema and API change coordination. A credible post-competition path can expand into regulated release assurance, data-platform change certification, cross-repository migration orchestration, agent capability certification, institutional memory governance, and enterprise-agent fleet control.
The product should remain focused on proof-carrying change rather than becoming a generic chatbot, generic workflow builder, or generic agent marketplace. (See Non-Goals and Red Lines for strict boundaries).
ChangeMesh source code is All Rights Reserved. Third-party dependencies retain their respective licenses. See LICENSE and docs/BUILD_PERIOD_DISCLOSURE.md.




