Skip to content

Latest commit

 

History

History
2399 lines (1719 loc) · 43.9 KB

File metadata and controls

2399 lines (1719 loc) · 43.9 KB

PROVENANCE Architecture — v0.2

Status

BOOTSTRAP MODULAR ARCHITECTURE
IMPLEMENTATION NOT YET FROZEN

This document defines the intended architecture of QSOLKCB/PROVENANCE.

PROVENANCE is designed as a modular evidence system:

small core
+
independent verification
+
optional storage
+
optional adapters
+
optional interfaces
+
optional presentation

The modules cooperate.

They must not become inseparable.

Normative engineering law remains in:

Project context and engineering ancestry remain in:


1. Architectural Objective

PROVENANCE is a framework-neutral system for recording, preserving, relating, and independently verifying evidence about actions performed by:

AI systems
agents
applications
tools
services
humans
automated workflows
distributed systems

Its primary optimization target is:

MINIMUM_OVERHEAD
subject to
ZERO_UNDECLARED_LOSS_OF_EVIDENTIARY_ACCURACY

The system should perform no work that is unnecessary to establish the declared evidence contract.

Performance may reduce:

latency
CPU time
memory
storage duplication
CI time
network traffic

It must not silently reduce:

evidence fidelity
verification strength
observation honesty
custody integrity

2. Architectural Style

PROVENANCE follows a modular architecture inspired by systems where a stable core is surrounded by replaceable engines, interfaces, hosts, and extensions.

Conceptually:

                     ┌─────────────────┐
                     │ provenance-ui   │
                     └────────┬────────┘
                              │
                    ┌─────────▼─────────┐
                    │ provenance-mcp    │
                    └─────────┬─────────┘
                              │
      ┌───────────────────────┼────────────────────────┐
      │                       │                        │
┌─────▼──────────┐    ┌───────▼────────┐      ┌────────▼────────┐
│ provenance-cli │    │ provenance-api │      │ provenance-     │
│                │    │ / embedding    │      │ adapters        │
└─────┬──────────┘    └───────┬────────┘      └────────┬────────┘
      │                       │                        │
      └───────────────────────┼────────────────────────┘
                              ▼
                     ┌─────────────────┐
                     │ provenance-core │
                     └────────┬────────┘
                              │
              ┌───────────────┼────────────────┐
              ▼               ▼                ▼
      provenance-store  provenance-verify  schemas/spec

No outer module may redefine core evidence semantics.


3. Monorepo First

PROVENANCE should initially remain one repository.

Module separation does not require separate GitHub repositories.

Current documentation-oriented repository structure:

PROVENANCE/
├── README.md
├── README4AIs.md
├── AGENTS.md
├── LICENSE
├── docs/
│   ├── README.md
│   ├── GETTING_STARTED.md
│   ├── INSTRUCTIONS.md
│   ├── ARCHITECTURE.md
│   ├── INVARIANTS.md
│   ├── ROADMAP.md
│   └── subsystem contracts...
├── provenance_core/
├── provenance_verify/
├── provenance_store/
├── provenance_custody/
├── provenance_export/
├── provenance_mcp/
├── provenance_cli/
├── provenance_ui/
├── provenance_adapters/
├── provenance-cli/
├── scripts/
└── tests/

The repository root intentionally keeps only project-entry and agent-guidance Markdown. Detailed specifications live under docs/.

Do not create empty modules merely to make the tree look mature.

A module should exist because it has a distinct contract.


4. Module Rule

Each module must satisfy:

ONE PRIMARY RESPONSIBILITY
CLEAR INPUT CONTRACT
CLEAR OUTPUT CONTRACT
NO HIDDEN AUTHORITY
REPLACEABLE WHERE PRACTICAL

Modules may depend downward.

Core modules must not depend upward.

Preferred direction:

UI
 ↓
MCP / CLI
 ↓
Adapters / Store
 ↓
Core
 ↓
Canonical Evidence Contract

Verification consumes evidence independently:

Evidence
   ↓
provenance-verify

The verifier must not require:

UI
MCP
Ollama
OpenAI
database server
specific adapter
original monitored application

to validate a finalized evidence bundle.


5. provenance-core

provenance-core is the smallest and most protected module.

It defines the universal evidence semantics.

Its responsibilities are limited to:

evidence classification
event model
artifact references
relationships
canonicalization
content identity
manifest structures
custody structures
failure/gap representation

It must remain:

framework-neutral
provider-neutral
storage-neutral
UI-neutral
transport-neutral
policy-neutral

It must not contain special knowledge of:

OpenAI
Claude
Gemini
Grok
Ollama
MCP
legal systems
medical systems
GitHub
databases
web browsers

Those belong above the core.


6. Core Data Types

The initial core should have very few universal concepts.

6.1 Artifact

An Artifact identifies retained content.

Examples:

prompt
response
document
image
tool payload
API response
source file
configuration
binary
log fragment

Conceptually:

Artifact
├── schema
├── content identity
├── structured record identity
├── byte count
├── media type
├── retention state
└── optional acquisition metadata

The architecture must distinguish:

CONTENT_RETAINED

from:

DIGEST_ONLY

A cryptographic digest is not the original artifact.


6.2 Event

An Event records an occurrence.

Conceptually:

Event
├── schema
├── event identity
├── evidence class
├── actor reference
├── operation
├── time information
├── input references
├── output references
├── relationships
├── observer
└── extensions

An event and an artifact are different things.

EVENT
answers:
"What occurred?"

ARTIFACT
answers:
"What exact content is this?"

6.3 Relationship

Relationships connect evidence without inventing meaning.

Examples:

previous
parent
triggered_by
input_to
output_of
derived_from
captured_from
approved_by
supersedes

Temporal adjacency does not automatically become causality.


6.4 Custody Event

Custody is append-only evidence describing handling.

Examples:

captured
stored
transferred
copied
exported
redacted
encrypted
verified
signed
released

Corrections append.

History does not rewrite.


6.5 Manifest

A Manifest identifies a collection of evidence.

Preferred shape:

ManifestCore
├── schema
├── event identities
├── artifact identities
├── custody identities
└── scope

        ↓ canonicalize + hash

ManifestEnvelope
├── core
└── core identity

The digest is computed over the core.

The envelope may then carry the digest.

No circular self-hashing.


7. Evidence Classification

The core evidence classes begin with:

OBSERVED
DECLARED
DERIVED

These describe provenance.

They are not truth scores.

OBSERVED != TRUE
DECLARED != FALSE
DERIVED != SPECULATIVE

Example:

An Ollama adapter directly receiving a response body may classify those received bytes as:

OBSERVED

The model name supplied by Ollama may be:

DECLARED

A SHA-256 digest calculated by PROVENANCE is:

DERIVED

8. Canonicalization

Canonicalization belongs in provenance-core.

It applies only to structured PROVENANCE records.

It must not modify raw evidence.

RAW ARTIFACT
     │
     ├── retained bytes
     └── content digest

STRUCTURED RECORD
     │
     ▼
canonical bytes
     │
     ▼
record digest

Target property:

same evidence values
+
same schema version
+
same canonicalization version
=
same canonical bytes

and:

same canonical bytes
+
same hash algorithm
=
same digest

9. Content Identity and Domain Separation

Raw artifact content uses ordinary algorithm-qualified cryptographic identity:

sha256:<SHA-256 of exact artifact bytes>

Raw artifact hashes are deliberately not domain-separated so that ordinary forensic and cryptographic tools can independently reproduce them.

Structured PROVENANCE record identities should use explicit semantic domains where appropriate:

PROVENANCE/ARTIFACT-RECORD/v1
PROVENANCE/EVENT/v1
PROVENANCE/MANIFEST/v1
PROVENANCE/CUSTODY/v1
PROVENANCE/CHECKPOINT/v1

This prevents identical structured canonical bytes used in different semantic roles from being accidentally treated as the same kind of record while preserving standard content identity for source artifacts.


10. provenance-verify

provenance-verify is an independent verification engine.

Its rule is:

RECOMPUTE
NOT TRUST

It consumes evidence.

It does not modify it.

The Phase 2 physical bundle contract is defined in BUNDLE.md. The bootstrap verifier validates canonical manifest/event/artifact records, exact physical membership, retained-content byte counts and SHA-256 identities, self-hash exclusion, and internal reference closure.

Typical verification operations include:

canonical form
content hashes
byte counts
manifest membership
event links
custody links
signatures
external anchors
bundle closure
schema validity

Output may include:

VERIFIED
FAILED
UNKNOWN
NOT_PRESENT
NOT_APPLICABLE

Verification must never perform:

repair
rewrite
normalization-in-place
historical correction

11. Verification Is Multidimensional

PROVENANCE must not collapse all assurance into one Boolean.

Example:

integrity = VERIFIED
schema = VERIFIED
custody = PARTIAL
signature = NOT_PRESENT
replay = NOT_POSSIBLE
anchor = NOT_PRESENT

This is more honest than:

trusted=true

or:

trusted=false

A valid historical record does not require every possible assurance dimension.


12. provenance-store

provenance-store handles persistence.

Storage is an implementation concern.

It must not define evidence meaning.

Possible backends include:

filesystem
content-addressed directory
SQLite
database
object storage
append-only service
remote custody store

The first backend should be simple.

A local filesystem or similarly lightweight store is preferred initially.

The Phase 3 reference backend is implemented in provenance_store and documented in STORE.md. It keeps mutable store state separate from verifier-compatible immutable snapshots. Content-addressed objects are published without overwrite; snapshots are independently verified before the mutable HEAD pointer advances.


13. Storage Model

The store should distinguish:

ARTIFACT CONTENT
ARTIFACT RECORDS
EVENT RECORDS
MANIFESTS / SNAPSHOTS
FUTURE CUSTODY RECORDS
DERIVED INDEXES

The Phase 3 reference store does not yet create custody records. Custody semantics belong to Phase 4.

Indexes are disposable.

Evidence is not.

If an index can be rebuilt, it should not become authoritative merely because querying it is faster.


14. Content Deduplication

Exact content may be stored once and referenced many times.

Example:

EVENT 1 ───┐
EVENT 2 ───┼──► sha256:X
EVENT 3 ───┘

Preferred:

store artifact X once
reference X three times

Deduplication must use strong content identity.

Never:

same filename
→ assume same artifact

15. provenance-adapters

Adapters connect monitored systems to the core.

Adapters are optional modules.

Potential adapters include:

ollama
openai
anthropic
gemini
grok
llama.cpp
generic HTTP
CLI/process
filesystem
agent framework
custom application

Each adapter must declare:

what it observes
what it does not observe
what values are declared externally
what values PROVENANCE derives
failure behavior

16. Adapter Rule

Adapters translate.

They do not invent.

Example:

provider did not expose immutable model revision

must remain:

immutable model revision = NOT_OBSERVED

not:

immutable model revision = guessed

Provider-specific metadata may be retained through extensions.

It must not reshape the universal core schema around one vendor.


16A. POSIX-First Local Portability

The local reference implementation should remain usable on ordinary Unix-like systems without requiring a desktop environment or network service.

Baseline principles:

terminal-native
POSIX-style filesystem semantics
POSIX advisory record locks
same-process mutexes around process-owned POSIX lock files
integer evidentiary time arithmetic
no Bash requirement
no systemd requirement
no GNU-command requirement
no Node requirement
no automatic network dependency

Host-specific tools are optional observers.

For example:

chronyc present
→ inspect existing host clock discipline

chronyc absent
→ use local system clock

explicit operator diagnostic
→ ntpdate -q <server>

PROVENANCE should execute subprocesses through explicit argv vectors, never by interpolating evidence into shell command strings.

Because classic POSIX record locks are process-owned, a record lock alone is not sufficient for multiple threads or instances in one process. Lock-file open/close operations must be serialized by the same process-local mutex used around acquisition of the POSIX record lock.

Presentation may use RFC3339/base-60 clock notation, but custody time calculations remain integer-only.


17. Ollama Reference Adapter

The first AI adapter is Ollama.

Reasons:

local
open integration surface
no cloud dependency
reproducible CI setup
easy request/response capture
small models available

The Phase 5 reference implementation is:

provenance_adapters.OllamaAdapter
adapter id = provenance-adapter:ollama/v1
transport = loopback HTTP only
endpoint = /api/generate
stream = false
dependencies = Python standard library + existing PROVENANCE modules

It retains the exact request bytes prepared by the adapter before transport and retains any HTTP response body bytes before parsing.

The transport boundary is enforced by an adapter-owned urllib opener with environment proxies disabled and redirects rejected. The documented localhost spelling is canonicalized to a literal loopback address to avoid DNS.

Prepared request evidence does not independently prove peer receipt. Transport and parse failures are recorded with COLLECTION_FAILED events and finalized evidence rather than silently disappearing.

Phase 5 uses one fresh evidence store/custody pair per exchange. Each unique evidence root is reserved with a same-process mutex plus a POSIX fcntl advisory record lock on an operational .ollama-observation.lock file. Roots are acquired in stable filesystem-identity order, and the reservation is held across freshness validation through complete observation/failure finalization. Freshness is re-read from disk after acquiring the reservation. This prevents concurrent threads, adapter instances, or cooperating processes that share either evidence root from both passing the freshness gate, and prevents deterministic event identities from collapsing repeated byte-identical calls until the core has an evidence-supported occurrence discriminator.

Adapter-retained HTTP payload bytes use one role-neutral artifact media type. Transport role is represented by events/custody so identical bytes can legitimately appear as both prepared request and observed response without rebinding content identity metadata.

Every finalized failure binds a canonical failure-detail artifact containing a stable adapter category, optional observed HTTP status, and diagnostic detail. Zero-length observed HTTP bodies remain distinct from no observed body.

The response model field is treated as a declaration by Ollama. It is not promoted into independently verified model provenance, and it is not used as the custody actor.

Capture times are expressed through custody observations rather than by adding provider-specific timestamp fields to the universal event core.

The Ollama adapter exists primarily to test the evidence architecture.

It must not define AI provenance semantics for every other provider.

Ollama's OpenAI-compatible endpoints are intentionally not used to redefine Phase 5. OpenAI-compatible HTTP remains a later generic interoperability surface behind the adapter boundary.


17A. Generic Adapter Interface

Phase 10 adds a provider-neutral adapter contract without changing the universal evidence schema.

The reference contract declares:

adapter identity
source kind
observation boundary
extension namespace

Adapter-specific execution produces directly captured payload artifacts plus conservatively DECLARED request/invocation descriptors and ordinary EventEnvelope values. Transport/provider metadata is retained as a separate DECLARED artifact rather than adding fields to EventCore.

The first two executable reference adapters are intentionally structurally different:

GenericHTTPAdapter
  request descriptor + request body
  response body / observed prefix
  HTTP transport boundary

ProcessAdapter
  argv descriptor + stdin
  stdout + stderr
  direct child-process boundary

Both flow through the same build_observation() and persist_observation() path.

Credentials are outside the ordinary evidence path unless the operator explicitly places them in captured content:

HTTP header values = runtime-only
inherited process environment = runtime-only

HTTP target URLs and process argv are evidence-bearing inputs. Operators must therefore avoid putting secrets in URLs/query strings or argv when those values should not be retained.

Failure does not erase observation. Failed HTTP/process operations return a persistable observation whose completion event is COLLECTION_FAILED.

See ADAPTERS.md.


18. Ollama CI Contract

The initial real-model lane uses a GitHub Actions matrix with separate clean runners. Each matrix job starts its own Ollama server and pulls one small reference model.

Current reference matrix:

qwen2.5:0.5b
qwen2:0.5b

The Ollama runtime is version-pinned in the workflow and its downloaded installer script is checksum-verified before execution.

A real inference test should verify the chain rather than exact generated language.

Example flow:

PROMPT
  ↓
OLLAMA REQUEST
  ↓
MODEL EXECUTION
  ↓
OLLAMA RESPONSE
  ↓
PROVENANCE EVENTS
  ↓
ARTIFACT HASHES
  ↓
MANIFEST
  ↓
VERIFY

Assertions should cover:

prepared request retained
successful or error response bytes retained when observed
failure evidence finalized on transport/parse failure
request identity recomputes
response identity recomputes
adapter identity recorded
model metadata classification correct
event relationships valid
manifest valid
verification passes

Then mutate retained evidence:

change one byte
↓
verification fails

Do not require:

prompt X always produces exact string Y

Model output is the observed evidence.

It is not the regression oracle.


19. provenance-mcp

provenance-mcp exposes PROVENANCE through the Model Context Protocol.

MCP is an interface.

It is not part of the evidence core.

Initial tools may include:

provenance.record
provenance.inspect
provenance.verify
provenance.finalize
provenance.export

Potential resources:

provenance://event/<id>
provenance://artifact/<identity>
provenance://manifest/<id>
provenance://custody/<id>
provenance://schema/<version>

20. MCP Evidence Boundary

An AI calling:

provenance.record(...)

does not make every supplied value independently observed.

If an AI says:

"I called tool X because Y"

through MCP, that information is normally:

DECLARED

unless PROVENANCE independently observed the relevant action or rationale.

The MCP server must not upgrade self-report into observation.


21. MCP Transport

Initial implementation should prefer the lowest-complexity useful transport.

Likely progression:

stdio
  ↓
local integration proven
  ↓
optional remote transport

Remote transport must not become mandatory for local evidence recording.

A local application should be able to use PROVENANCE without operating a network service.


22. provenance-cli

The CLI is the simplest human and automation interface.

The preferred interactive implementation direction is a Rust TUI with keyboard-first slash-command discovery. Typing / should open/filter a compact command palette rather than requiring users to memorize flags for common interactive operations.

Initial commands may conceptually include:

provenance record
provenance verify
provenance inspect
provenance finalize
provenance export

The CLI should expose core functionality directly.

It should not contain a second implementation of verification semantics.

Provider authentication belongs behind an interface boundary. Supported modes may include API keys, OAuth device authorization, OAuth browser/loopback authorization, or no authentication for local endpoints. Tokens and credentials are operational secrets, not ordinary evidence payloads.

OpenAI-compatible API surfaces may be supported as a generic interoperability adapter because many providers and local/open-source systems implement similar request/response shapes. That compatibility must remain outside the universal evidence core.


23. provenance-ui

The UI is an evidence viewer.

It has no evidentiary authority.

The initial viewer should be served by a tiny local HTTP server bound to loopback by default and implemented with pure HTML/CSS plus minimal vanilla JavaScript. No frontend framework or Node runtime is required unless a later concrete requirement earns that complexity.

Its job is to make the chain understandable.

Primary views should include:

timeline
event graph
artifact inspector
custody history
verification status
evidence gaps

The UI consumes existing evidence.

It does not define evidence.

Network exposure is explicit:

default = 127.0.0.1
LAN/public = operator opt-in

Do not expose unrelated inetd-style utility services as part of the viewer.


24. UI Evidence Graph

The central visual model should be the evidence chain.

Example:

USER INPUT
    │
    ▼
PROMPT ARTIFACT
sha256:...
    │
    ▼
MODEL INVOCATION
    │
    ├────► TOOL CALL
    │         │
    │         ▼
    │      TOOL RESULT
    │
    ▼
MODEL RESPONSE
sha256:...
    │
    ▼
APPLICATION ACTION
    │
    ▼
HUMAN REVIEW

Selecting a node should expose:

WHO
WHAT
WHEN
WHERE
HOW
WHY-EVIDENCE
CLASSIFICATION
IDENTITY
CUSTODY
VERIFICATION

25. Evidence Gaps in the UI

Evidence gaps must be visually explicit.

Never quietly connect:

EVENT 15
   ↓
EVENT 16

when collection was interrupted.

Instead:

EVENT 15
   │
   ▼
██████████████████████
█  EVIDENCE GAP      █
█ collection failed █
██████████████████████
   │
   ▼
EVENT 16

A missing record is not evidence that nothing happened.


26. UI Technology

The initial UI should remain lightweight.

Preferred direction:

HTML
CSS
minimal JavaScript

unless requirements later justify something heavier.

Avoid introducing a large application framework merely to display evidence graphs and metadata.

The UI should be replaceable without affecting:

recording
storage
verification
evidence identity

26A. Portable Forensic Package

Phase 11 adds an archival producer/verifier pair outside the evidence core.

finalized provenance.bundle.v1
stable custody snapshot
schema/version metadata
recomputed verification metadata
recomputed declared gaps
        ↓
provenance.forensic-package.v1

The package envelope binds every member by portable relative path, SHA-256 content identity, and byte count.

The embedded Phase 2 evidence bundle remains unchanged and is independently re-verified after copying.

Package finalization and evidence collection scope are independent dimensions:

package_state = FINALIZED
evidence_scope = open | closed

The producer uses hidden staging plus independent verification before atomic publication. The verifier requires exact physical membership and rejects undeclared directories as well as undeclared files and unsafe filesystem objects.

CLI and MCP call the same producer contract through provenance package and provenance.package. Their older snapshot-copy export interface remains unchanged.

See PACKAGE.md.


27. Observation Topologies

Adapters may observe systems through several topologies.

Native Hook

APPLICATION
    ├── normal operation
    └── provenance observation

Middleware / Proxy

APPLICATION
    ↓
PROXY
    ↓
EXTERNAL SYSTEM

Sidecar

APPLICATION ─────► normal system
     │
     └───────────► PROVENANCE

External Observer

SYSTEM
   ↓
observable external effects
   ↓
PROVENANCE

Each has a different observation boundary.

PROVENANCE must preserve that distinction.


28. Non-Interference

No PROVENANCE module may silently alter monitored behavior.

Forbidden hidden operations include:

rewrite prompt
rewrite response
change model settings
retry request
suppress tool call
change tool output
change decision
reorder monitored action

An integrating application may explicitly choose such behavior.

That application decision then belongs to the monitored system, not PROVENANCE.


29. Hot-Path Architecture

The monitored hot path should perform the minimum work required for accurate capture.

Preferred:

observe
→ capture bytes/reference
→ minimal metadata
→ enqueue/store
→ return

Move nonessential work off the hot path:

indexing
search
timeline construction
visualization
report generation
derived analysis

Cryptographic operations may be streamed or deferred only when doing so does not create an undeclared integrity gap.


29A. Phase 13 Exact Parallel Verification

Performance hardening may parallelize independent physical verification work only when observable verifier semantics remain equal to the serial reference.

Phase 13 applies:

bounded worker execution
+ deterministic input-order reduction
+ serial reference parity

to artifact/content checks, event checks, and forensic-package member hashing.

Worker completion order is never report order.

The required invariant is:

optimized VerificationReport
==
serial-reference VerificationReport

including check order and error order.

Parallel scheduling is bounded and batch-limited; configured worker count is not treated as evidence of actual overlap. CI separately witnesses overlapping work and enforces the live-worker cap.

Exact parity is claimed for stable evidence inputs. Verification still does not provide an atomic snapshot of a concurrently mutated external directory tree.

See PERFORMANCE.md.


30. Bounded Collection

PROVENANCE must not use unbounded memory merely to avoid admitting that evidence was lost.

Collectors need explicit bounded behavior.

Possible outcomes:

RECORDED
PARTIALLY_RECORDED
DROPPED
COLLECTION_FAILED
EVIDENCE_GAP_OPENED

Silently dropping evidence while presenting continuity is forbidden.


31. Backpressure

PROVENANCE should not secretly become application flow control.

Possible integration policies include:

best effort
bounded buffer
synchronous evidence capture
application-defined fail closed
application-defined fail open

The host chooses the policy.

PROVENANCE records or exposes which policy applies.


32. Failure Architecture

Failure is evidence.

Examples:

capture failure
serialization failure
storage failure
hash failure
queue overflow
network failure
permission failure
unsupported field
process termination

Correct behavior:

DETECT
  ↓
PRESERVE WHAT IS KNOWN
  ↓
MARK UNKNOWN / MISSING PORTION
  ↓
REPORT GAP

Never:

FAILURE
  ↓
INVENT REPLACEMENT

33. Time and Ordering

PROVENANCE should retain both time and ordering information without treating them as identical.

Possible fields:

capture_time
declared_event_time
monotonic_time
sequence
clock_source
clock_precision

Do not manufacture a global total order when evidence supports only partial ordering.


34. Replay

Replay is optional evidence.

It is not a universal validity requirement.

HISTORICAL OBSERVATION
!=
LATER REPLAY

Replay may be:

VERIFIED
FAILED
NOT_POSSIBLE
NOT_ATTEMPTED

The inability to replay an old model or external system does not invalidate authentic historical evidence.


35. Reference and Optimized Paths

Where practical, important operations should retain a straightforward reference path.

Example:

                 ┌── reference verifier
evidence ────────┤
                 └── optimized verifier

results must agree

Allowed optimization mechanisms include:

streaming
deduplication
incremental verification
proven-equivalent caching
bounded parallelism
verified dependency reuse

Required condition:

OPTIMIZED SEMANTICS
==
REFERENCE SEMANTICS

36. Caching

Cache entries are not trusted merely because they exist.

Reuse requires:

COMPLETE EFFECTIVE INPUT IDENTITY
+
VALIDATED OUTPUT IDENTITY
+
SUCCESSFUL PRIOR GENERATION
+
PRESERVED SEMANTIC CONTRACT

A failed or interrupted operation must never publish authoritative reusable state.

For cheap cryptographic verification:

prefer recomputation

over complicated caching.


37. Module Independence Test

Each optional module should be removable.

Deleting:

provenance-ui

must not invalidate evidence.

Deleting:

provenance-mcp

must not invalidate evidence.

Deleting:

provenance-adapters/ollama

must not break the verifier.

Replacing:

provenance-store

must not change core evidence semantics.

This is a central architectural test.


38. Interface Consistency

All interfaces should ultimately use the same core contracts.

CLI ─────┐
MCP ─────┤
UI ──────┤
SDK ─────┤
Adapters ┘
         ↓
   provenance-core

There must not be:

MCP evidence format
UI evidence format
CLI evidence format

with subtly different meanings.

One evidence contract.

Many interfaces.


39. Privacy

Evidence completeness does not require collecting everything.

An adapter should distinguish:

required evidence
optional context
unnecessary sensitive data
secret material

Credentials and private keys must not become ordinary evidence payloads.

Supported strategies may include:

digest-only retention
encryption
redaction derivative
content omission
access-controlled storage
retention policies

If content is omitted, that fact must remain explicit.


39A. Phase 14 Selective Disclosure

The reference privacy path creates a new derivative rather than mutating evidence:

verified retained source artifact
        ↓ deterministic redaction
new DERIVED artifact
        ↓
selective-disclosure package

The disclosure package withholds original source bytes while retaining the source content identity and exact original ArtifactRecord metadata as a witness.

Original retention and disclosure retention are separate concepts:

source package: CONTENT_RETAINED
selective disclosure: DIGEST_ONLY source witness
derivative: CONTENT_RETAINED

Standalone disclosure verification proves disclosed content integrity and DERIVED lineage. Exact transform verification requires the original source package and is reported separately.

The source package is held open and verified by descriptor before source bytes are used for derivation.

See PRIVACY.md.


40. Framework Neutrality

The architecture must support:

OpenAI
Anthropic
Google
xAI
Ollama
llama.cpp
future unknown providers
non-AI systems

without changing the universal evidence semantics.

Correct:

provider
   ↓
adapter
   ↓
PROVENANCE CORE

Wrong:

PROVENANCE CORE
=
one provider schema
+
patches for everyone else

41. Language Neutrality

The reference implementation may initially use one language.

The evidence contract must not depend on language-specific object representation.

Forbidden protocol authority includes:

Python repr()
pickle
process object IDs
dict insertion accidents
runtime hash()

Independent implementations must eventually be possible.


42. Dependency Direction

Hard direction:

provenance-ui
      ↓
provenance-mcp / provenance-cli
      ↓
provenance-adapters / provenance-store
      ↓
provenance-core

Independent branch:

evidence
   ↓
provenance-verify
   ↓
verification report

Forbidden dependency direction:

provenance-core
      ↓
provenance-ui

or:

provenance-verify
      ↓
Ollama adapter

43. Bootstrap Module Order

Do not build every module immediately.

Preferred implementation order:

PHASE 1
provenance-core
    canonical bytes
    artifact identity
    event
    manifest

PHASE 2
provenance-verify
    manifest verification
    tamper detection
    fixtures

PHASE 3
provenance-store
    simple local storage
    content deduplication

PHASE 4
provenance-adapters/ollama
    real AI observation
    GitHub Actions integration

PHASE 5
provenance-mcp
    stdio tools/resources

PHASE 6
provenance-cli
    inspection and verification

PHASE 7
provenance-ui
    evidence graph
    timeline
    gaps
    artifact inspection

PHASE 8
additional adapters
    OpenAI
    Claude
    Gemini
    Grok
    generic HTTP

Signatures, external anchoring, distributed custody, and formal verification come later unless a concrete requirement pulls them forward.


44. CI Architecture

CI should also remain modular.

core.yml
    canonicalization
    identity
    manifests
    tamper fixtures

verify.yml
    verifier contract
    malformed evidence
    integrity failures

ollama.yml
    real local model smoke test

mcp.yml
    MCP protocol surface

ui.yml
    lightweight presentation tests

full.yml
    integration / release-grade validation

Path filtering or equivalent selective execution may reduce routine CI cost once module boundaries stabilize.

Invariant-sensitive changes must still escalate appropriately.


45. Ollama CI Principle

The Ollama test validates PROVENANCE.

It does not validate Ollama's intelligence.

Do not assert:

PROMPT X
→ EXACT RESPONSE Y

Assert:

REQUEST OBSERVED
RESPONSE OBSERVED
ARTIFACT IDENTITIES VALID
EVENT RELATIONSHIPS VALID
MANIFEST VALID
VERIFY PASS

Then:

TAMPER WITH ONE BYTE
→ VERIFY FAIL

46. MCP Self-Demonstration

A useful integration test is:

AI / client
   ↓
calls PROVENANCE through MCP
   ↓
PROVENANCE records interaction
   ↓
client requests evidence chain
   ↓
PROVENANCE returns verifiable record

This demonstrates the system using its own public integration surface.

Self-observation must still preserve evidence classification boundaries.


46A. Optional Trust Records

Phase 12 adds provenance_trust as an optional producer of detached authenticity records.

finalized forensic package
        ↓
optional provenance-trust producer
        ↓
signature / external-anchor sidecars

package + sidecars
        ↓
provenance-verify
        ↓
integrity / signature / external-anchor dimensions

The package identity does not change when trust records are added.

Reference mechanisms are OpenSSH Ed25519 SSHSIG signatures over exact package.json bytes and Git commit anchors containing an exact canonical package-identity payload.

Private signing keys are producer inputs only. They do not enter the evidence core, package, verifier, or MCP protocol.

External-anchor mechanisms are replaceable and optional. Phase 12 performs no automatic network access and does not require blockchain infrastructure.

See TRUST.md.


47. Deployment Model

The same modules should support several deployment sizes.

Minimal

application
+
provenance-core
+
local store

Developer / AI

application
+
adapter
+
core
+
store
+
MCP

Investigator

evidence bundle
+
provenance-verify
+
CLI/UI

Large deployment

many adapters
+
core-compatible evidence service
+
remote storage
+
MCP/API
+
UI
+
external anchors

Scale changes.

Evidence semantics do not.


48. Architectural Anti-Patterns

Avoid making any of the following mandatory:

central server
database cluster
cloud account
AI provider
MCP
UI
blockchain
message broker
container platform
replay
formal prover
JavaScript framework

Optional infrastructure must remain optional.


49. No Mandatory Blockchain

Chain of custody is not synonymous with blockchain.

External anchoring may use:

digital signature
Git commit
signed release
timestamp service
transparency log
DOI record
append-only service
blockchain

PROVENANCE defines the evidence.

Anchoring mechanisms are replaceable.


50. Architectural Success Tests

The architecture succeeds if all of these are possible:

Tiny integration

small CLI
+
provenance-core
+
local files

Local AI integration

Ollama
+
adapter
+
core
+
store

MCP integration

AI client
+
provenance-mcp
+
core

Human investigation

old evidence bundle
+
provenance-verify
+
UI

Large platform

many systems
+
many adapters
+
shared store
+
same evidence contract

51. Five-Year Test

Assume five years have passed.

The monitored application no longer exists.

The AI model is unavailable.

The original developers are gone.

The UI has been completely rewritten.

The MCP protocol implementation has changed.

An investigator still possesses:

artifacts
events
custody records
manifests
schemas
public specifications

A current independent verifier should still be able to determine:

what was captured
what was retained
what was declared
what was derived
which bytes were hashed
which identities recompute
where gaps exist
which relationships are supported
what changed
what verifies
what cannot be known

If preserving evidence requires resurrecting the old UI, MCP server, Ollama adapter, or original application:

ARCHITECTURE FAILED

Final Architecture Law

PROVENANCE should resemble a collection of small cooperating instruments:

CORE
VERIFY
STORE
ADAPTERS
MCP
CLI
UI

rather than one monolithic application.

Each module should know only what it needs to know.

Each optional surface should remain replaceable.

The core must remain boring.

That is a feature.

SMALL CORE
CLEAR MODULES
ONE EVIDENCE CONTRACT
OPTIONAL INTEGRATIONS
INDEPENDENT VERIFICATION
MINIMUM OBSERVER EFFECT

Architectural checksum:

CAPTURE ONLY WHAT THE CONTRACT REQUIRES.

PRESERVE EXACTLY WHAT WAS CAPTURED.

KEEP THE CORE INDEPENDENT OF THE INTERFACES.

MAKE EVERY OPTIONAL MODULE REPLACEABLE.

MAKE THE EVIDENCE SURVIVE THE SOFTWARE THAT CREATED IT.

GET OUT OF THE MONITORED SYSTEM'S WAY.

Phase 15 — Distributed Custody Architecture

The reference distributed path is a signed store-and-forward handoff, not a global ledger.

source forensic package
  ├── source evidence identity
  └── sender custody snapshot
          ↓
signed transfer offer
          ↓
transfer bundle
          ↓ offline / delayed transport
receiver verifies package + offer
          ↓
receiver-local CAPTURED/STORED/VERIFIED chain
          ↓
signed receipt

Dependency direction remains optional and outward:

provenance-transfer
    ↓
provenance-export / provenance-custody / provenance-trust

transfer evidence
    ↓
provenance-verify

No cross-system timestamp comparison creates ordering authority. The verifier reports explicit causal edges and PARTIAL ordering only.

See TRANSFER.md.


Phase 16 — Release-Grade Trust Architecture

Phase 16 adds a validation layer around the existing modules without changing evidence semantics.

focused module CI
      ↓
routine development feedback

exact release candidate ref
      ↓
full.yml
  ├── pinned Python + Rust
  ├── complete invariant/tamper/integration corpus
  ├── CLI build/tests
  ├── canonicalization cross-seed recheck
  ├── clean-tree check
  └── real Ollama matrix
      ↓
release-gate
      ↓
eligible for Phase 17 freeze

The release lane composes the existing module contracts. It does not define a new evidence format, verifier, adapter, package, or custody model.

The high-assurance lane deliberately avoids authoritative dependency caches and uses fresh checkouts. Toolchain/action versions are pinned at the workflow level where practical. The managed GitHub runner family remains an execution environment rather than evidence authority.

A successful run binds assurance to the exact tested repository commit. Any implementation change requires a new Phase 16 run before Phase 17 freeze.

Formal proof remains downstream:

Phase 16 engineering gate
        ↓
Phase 17 immutable implementation freeze
        ↓
Phase 18 Lean formalization + archival release

See RELEASE.md.


Phase 18 — Formal Verification and Archival Architecture

Formal verification is downstream of the immutable implementation target.

v1.0.0
0b1a2eea6c3c2b40a7f2a390fcd3410c75fab742
        ↓
explicit model/runtime bridge
        ↓
Lean 4.34.1 proof model
        ↓
independent proof checks
        ↓
formal evidence manifest
        ↓
archive bundle
        ↓
final archival tag + Zenodo DOI

The dependency direction is one-way:

frozen implementation
        ↓
formal model

NEVER

formal proof convenience
        ↓
silent frozen implementation rewrite

The proof source is not part of provenance-core and does not change evidence semantics.

Phase 18 currently formalizes only:

FV-01 self-hash exclusion
FV-02 append-only history extension
FV-03 classification non-promotion
FV-04 presentation non-interference

Runtime correspondence is explicit in FORMAL_VERIFICATION.md. The archival/DOI process is defined in ARCHIVAL_RELEASE.md.

If the formalization discovers an implementation defect requiring contract changes, v1.0.0 remains immutable and a new release candidate must be established.