Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
f0251b7
Phase 0: establish foundational research contract
EmergentMonk Oct 6, 2026
d833e29
Fix Phase 0 review findings
EmergentMonk Oct 6, 2026
9ea5418
Harden terminology heading validation
EmergentMonk Oct 6, 2026
7339c5f
Ignore fenced Markdown headings in Phase 0 validator
EmergentMonk Oct 6, 2026
57e8517
Add fenced-heading validator regressions
EmergentMonk Oct 6, 2026
8fd0699
Fix Codex Phase 0 review findings
EmergentMonk Oct 6, 2026
1e1fa9c
Fix latest Phase 0 review findings
EmergentMonk Oct 6, 2026
500f843
Harden Markdown blocks and experimental controls
EmergentMonk Oct 6, 2026
024897e
Fix raw HTML parser regex escaping
EmergentMonk Oct 6, 2026
20d12cc
Correct raw HTML parser regexes
EmergentMonk Oct 6, 2026
bf0bf6d
Fix parser edge cases and hypothesis identification
EmergentMonk Oct 6, 2026
0892df0
Fix remaining Markdown parser edge cases
EmergentMonk Oct 6, 2026
425892e
Fix Markdown inline and container edge cases
EmergentMonk Oct 6, 2026
0275331
Fix thematic break and ordered-list regexes
EmergentMonk Oct 6, 2026
266a68e
Fix list, prose, setext, and HTML parsing cases
EmergentMonk Oct 6, 2026
184831d
Fix hidden link-reference prose filtering
EmergentMonk Oct 6, 2026
6a83959
Fix multiline link-reference prose handling
EmergentMonk Oct 6, 2026
534666d
Fix reference metadata edge cases
EmergentMonk Oct 6, 2026
8e04372
Fix reference destination backslash literal
EmergentMonk Oct 6, 2026
7fd4593
Fix reference syntax and multiline title cases
EmergentMonk Oct 6, 2026
f0ae1f5
Fix latest Phase 0 Markdown edge cases
EmergentMonk Oct 6, 2026
f1bc596
Fix list prose and Unicode whitespace cases
EmergentMonk Oct 6, 2026
689c0a2
Fix nested list prose and size validation
EmergentMonk Oct 6, 2026
1c63abb
Fix nested list and list-reference parsing
EmergentMonk Oct 6, 2026
f1cdd8d
Unify Markdown container parsing
EmergentMonk Oct 6, 2026
1d6c34e
Add Markdown parser regression matrix
EmergentMonk Oct 6, 2026
659640e
Harden Markdown regression coverage
EmergentMonk Oct 6, 2026
d6d1849
Fix visual tab padding regression
EmergentMonk Oct 6, 2026
c775a08
Harden container and inline Markdown parsing
EmergentMonk Oct 6, 2026
7b24351
Preserve list paragraph state in visibility pass
EmergentMonk Oct 6, 2026
cc90829
Distinguish list content from rendered paragraph text
EmergentMonk Oct 6, 2026
b0e1e79
Fix Unicode-only list paragraph state
EmergentMonk Oct 6, 2026
cb405dd
Fix lazy list paragraph state in containers
EmergentMonk Oct 6, 2026
2708802
Strengthen paragraph container regressions
EmergentMonk Oct 6, 2026
b8aacc5
Track paragraph quote depth in visibility records
EmergentMonk Oct 6, 2026
3ea0ba8
Harden container-aware rendered Markdown
EmergentMonk Oct 6, 2026
75c74db
Strengthen rendered Markdown edge regressions
EmergentMonk Oct 6, 2026
778105b
Harden quote fence and lazy marker regressions
EmergentMonk Oct 6, 2026
f5b9b2f
Harden CommonMark rendering regressions
EmergentMonk Oct 6, 2026
44c9e0c
Unify Markdown ownership regression handling
EmergentMonk Oct 7, 2026
c0946e4
Replace divergent Markdown scans with one pinned CommonMark parse
EmergentMonk Oct 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions .github/workflows/phase0.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
name: Phase 0 Contract

on:
push:
branches: [main]
pull_request:

permissions:
contents: read

jobs:
validate:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python: ["3.12", "3.14"]
steps:
- name: Check out repository
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python }}
cache: pip

- name: Install pinned validator dependencies
run: python3 -m pip install --require-hashes -r requirements.txt

- name: Validate research contract
run: python3 scripts/validate_phase0.py

- name: Run tests
run: python3 -m unittest discover -s tests -v
62 changes: 62 additions & 0 deletions HYPOTHESES.md
Original file line number Diff line number Diff line change
@@ -1 +1,63 @@
# Hypotheses

CONSTRAINT-SHIFT begins with six falsifiable hypotheses. A hypothesis may be supported, weakened, rejected, split, or replaced by evidence. It must not be silently reworded to fit results.

## H1 — Specification Primacy

**Claim:** As independently measured AI implementation capability increases, the relative benefit of persistent specifications and machine-checkable contracts over transient prompt-only workflows increases for preserving, regenerating, and changing intended software behaviour.

**Candidate measurements:** predeclared AI capability strata measured independently of H1 outcomes; regeneration success against a fixed acceptance suite; change-propagation success; human implementation editing after regeneration; behavioural divergence across repeated implementations; and the interaction between capability level and workflow condition.

**Falsification pressure:** H1 is weakened if the persistent-specification advantage is flat, decreases, or disappears as independently measured implementation capability increases under matched tasks and resource budgets. A benefit observed at only one capability level supports specification utility under that condition but does not by itself support H1's directional claim.

## H2 — Verification Selection

**Claim:** AI-assisted development disproportionately benefits languages and toolchains that provide precise, machine-actionable diagnostics and enforceable constraints.

**Candidate measurements:** attempts to first successful compile; repair iterations; diagnostic-to-fix conversion; human interventions; static-analysis defects; token and wall-clock cost.

**Falsification pressure:** H2 is weakened if stronger machine-checkable constraints and diagnostics provide no reproducible improvement after controlling for model familiarity, ecosystem maturity, and task suitability.

## H3 — Legacy Preservation Paradox

**Claim:** AI assistance can extend the operational lifetime of legacy languages and systems by reducing maintenance cost associated with scarce human expertise.

**Candidate measurements:** predeclared expertise/scarcity strata measured independently of H3 outcomes; matched AI-assisted versus unassisted conventional-maintenance outcomes within each stratum; scarce-expert consultation hours consumed; defect-localization and repair success; active human effort; verification burden; behavioural regressions; and a predeclared maintenance-viability horizon or equivalent lifecycle proxy under a fixed cumulative maintenance budget.

**Falsification pressure:** H3 is weakened if AI assistance does not reduce maintenance burden relative to the matched unassisted baseline under scarce-expertise conditions, if the benefit does not persist or increase as expert access becomes scarcer, if the predeclared maintenance-viability horizon is not extended, or if apparent gains are offset by verification burden or behavioural risk. Short experiments may support only the declared lifecycle proxy, not literal calendar-year lifetime claims.

## H4 — Modding Mutation

**Claim:** AI-assisted development reduces the technical barrier to software modification and increases the number and diversity of executable modification attempts under a fixed effort budget.

**Candidate measurements:** executable modifications per unit time; human technical actions; distinct completed changes; failure rate; behavioural diversity.

**Falsification pressure:** H4 is weakened if AI-assisted workflows do not increase executable modification throughput or diversity once setup, debugging, and correction costs are included.

## H5 — Recombination Acceleration

**Claim:** When implementation cost is reduced while design intent and quality criteria are held fixed, the rate at which mechanics, systems, genres, and implementation patterns are recombined into executable prototypes increases under a fixed total effort budget.

**Candidate measurements:** a predeclared outcome-independent implementation-input measure recorded per attempt, such as active human time, wall-clock time, tool/agent steps, or reliable compute/token expenditure; successful executable recombinations per fixed total effort budget; retained source concepts; integration defects; and behavioural verification. Derived efficiency metrics such as cost per accepted prototype may be secondary outcomes but cannot establish the H5 mechanism.

**Falsification pressure:** H5 is weakened if an implementation-focused intervention fails to reduce the predeclared implementation-cost measure, if measured cost reduction does not increase executable recombination rate under the fixed total budget, or if recombination increases only when design ideation changes while implementation cost remains unchanged.

## H6 — Tool Legitimacy Gap

**Claim:** For otherwise equivalent software artifacts, disclosure of AI participation can alter perceived legitimacy independently of demonstrated artifact quality.

**Candidate measurements:** perceived quality, creativity, originality, authenticity, technical competence, trust, and willingness to use or recommend.

**Falsification pressure:** H6 is weakened if disclosure produces no reproducible difference, or if observed differences are explained by uncontrolled artifact, wording, sampling, or expectancy effects.

## Status language

Results should use restrained labels:

- **untested**
- **inconclusive**
- **supported under tested conditions**
- **weakened**
- **falsified under tested conditions**

“Proven” is not the default label for an empirical result.
44 changes: 44 additions & 0 deletions INVARIANTS.md
Original file line number Diff line number Diff line change
@@ -1 +1,45 @@
# Research Invariants

These rules constrain every experimental phase unless a later change explicitly amends the research contract and documents compatibility with prior evidence.

## I1 — No conclusion by model assertion
An LLM output, prediction, ranking, or explanation is not evidence for a project hypothesis by itself.

## I2 — Equivalent-task comparisons
Cross-language, cross-tool, and cross-agent comparisons must begin from an equivalent task contract. Language-specific accommodations must be declared rather than hidden.

## I3 — Preserve failures
Failed generations, compiler failures, abandoned repair paths, timeouts, and invalid outputs are part of the evidence and must not be silently discarded.

## I4 — Record intervention
Human intervention must be recorded at a useful granularity. Manual fixes may not be attributed to an agent.

## I5 — Record environment
Experiments must retain enough environment information to interpret the result, including model/tool identity, toolchain versions, task/specification identity, and execution platform where relevant.

## I6 — Separate generation from verification
A system that generates an implementation must not be treated as independent verification merely because it also says the implementation is correct.

## I7 — Observable contracts before equivalence claims
Implementation fungibility or behavioural equivalence claims require an explicit observable contract and a declared comparison procedure.

## I8 — No predetermined language winner
The project must not choose metric weights or task suites solely to produce a preferred language ranking.

## I9 — Negative results are publishable results
A result that contradicts the central thesis remains valid project output if the method and evidence are sound.

## I10 — Social experiments control the artifact
Tool-legitimacy experiments must hold the evaluated artifact constant across disclosure conditions unless artifact variation is itself the declared independent variable.

## I11 — Human-subject safeguards
Before recruitment or participant data collection begins, human-participant experiments must have a documented protocol covering informed consent, privacy and data minimization, data handling and retention, and any applicable ethics or institutional review. No participant data may be collected until required approvals are in place and the consent process is ready for use; if formal review is not required, that determination should be documented before recruitment.

## I12 — Motivation is not evidence
Historical analogies, anecdotes, popularity trends, and project origin stories may motivate a hypothesis but do not count as experimental confirmation.

## I13 — Reproducible validators
Machine-derived claims should be accompanied by reproducible validators or analysis code where practical.

## I14 — Contract changes are explicit
Changes to hypotheses, terminology, methodology, or invariants that affect interpretation of existing results must be versioned and explained.
73 changes: 73 additions & 0 deletions METHODOLOGY.md
Original file line number Diff line number Diff line change
@@ -1 +1,74 @@
# Methodology

## Purpose
CONSTRAINT-SHIFT is designed as an empirical research repository, not a collection of AI-development predictions. Each experiment should connect a falsifiable hypothesis to an operational definition, controlled procedure, retained evidence, and bounded conclusion.

## Common experiment lifecycle

1. **Select hypothesis.** Identify the exact hypothesis and sub-claim under test.
2. **Freeze task and analysis contract.** Before outcome inspection, define inputs, outputs, acceptance conditions, resource limits, stopping rules, inclusion/exclusion criteria, invalid-trial and timeout rules, missing-data handling, and the denominators used for primary rates.
3. **Declare factors.** Record independent variables such as language, compiler, agent, disclosure condition, or assistance mode.
4. **Declare controls.** Identify variables held constant and unavoidable differences.
5. **Run trials.** Preserve successful and failed trials.
6. **Verify independently.** Use compilers, tests, static checks, formal tools, or blinded evaluation appropriate to the claim.
7. **Retain evidence.** Store machine-readable results plus enough provenance to reproduce interpretation.
8. **Analyze.** Report distributions and uncertainty rather than only best-case examples.
9. **Bound the conclusion.** State what was tested and what was not.
10. **Attempt replication.** Prefer conclusions that survive reruns, alternative tasks, and changed models/toolchains.

## Experimental unit
The experimental unit must be explicit: one generation attempt, repair trajectory, task-language pair, legacy-maintenance task, modding task, participant evaluation, or another declared unit.

Repeated attempts from the same underlying task are not automatically independent observations.

## Comparability
Cross-language experiments must distinguish same specification from same implementation strategy, language-required boilerplate from task logic, ecosystem effects from core language/toolchain effects, model familiarity from language semantics, and compile-time failure from behavioural failure.

## AI execution metadata
Where available and permitted, record model/provider identifier, date, agent/tool version, relevant instructions, iteration count, compiler/tool cycles, reliable token/cost measures, human intervention, and termination reason.

If exact parameters are unavailable, record that limitation rather than inventing them.

## Toolchain metadata
Record language version, compiler/interpreter version, build flags, dependency lock state, analysis tools, operating system/architecture, test command, and materially relevant hardware.

## Evidence classes

### E0 — Motivation
Anecdote, historical example, external observation, or qualitative rationale. Useful for choosing a question; not confirmation.

### E1 — Single controlled trial
A reproducible result under one task/environment condition.

### E2 — Replicated controlled evidence
Repeated results under the same contract with retained failures and stable analysis.

### E3 — Cross-condition evidence
Results reproduced across materially different tasks, models, toolchains, systems, or participant samples.

## Metrics
Metrics are hypothesis-specific. Phase 0 approves categories, not fixed weights: success/failure, compile attempts, repair iterations, test-pass rate, static/formal outcomes, wall-clock duration, reliable token/cost data, human interventions, behavioural divergence, modification throughput, recombination completion, and controlled perception ratings.

Any composite score must publish its formula, normalization, weights, missing-data policy, and sensitivity to alternative weights.

## Statistical reporting
Inclusion, exclusion, invalid-trial, timeout, missing-data, and primary-denominator rules must be frozen before outcome inspection. Failures retained under I3 must not be reclassified after results are known merely to improve a reported rate.

When sample size permits, report sample count, distributions, uncertainty, exploratory versus confirmatory status, all exclusions and invalid trials with reasons, missingness, and dependence between repeated attempts.

Any deviation from the predeclared trial-handling or missing-data rules must be identified explicitly, justified, and accompanied by a sensitivity analysis showing the result under the original rule where technically possible.

## Human evaluation
Before recruitment or data collection, a human-participant protocol must define informed-consent procedures, privacy and data-minimization safeguards, data handling and retention, primary outcomes, and any applicable ethics or institutional review. Required approvals must be in place before recruitment or collection begins; if formal review is not required, document that determination beforehand.

Once those prerequisites are satisfied, studies should randomize or counterbalance where appropriate and hold the artifact constant when testing disclosure.

## Legacy-code safety
Legacy experiments should use public, synthetic, redistributable, or otherwise authorized code. Toy experiments cannot establish safety for production banking, trading, medical, industrial, or other critical systems.

## Reproducibility target
A retained experiment should make it possible for an independent operator to answer:

> What was attempted, under which contract, with which tools, what happened, and how was the result judged?

If those questions cannot be answered, the experiment is incomplete.
100 changes: 99 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
@@ -1,2 +1,100 @@
# CONSTRAINT-SHIFT
Research and experiments on how AI shifts software development from code authorship toward specification, constraints, verification, and machine-generated implementation.

**Research and experiments on how AI shifts software development from code authorship toward specification, constraints, verification, and machine-generated implementation.**

CONSTRAINT-SHIFT treats a simple observation as a research problem: when implementation becomes cheap to generate, the scarce work may move upward into defining intent, constraining admissible solutions, verifying outcomes, and selecting among candidate implementations.

## Central thesis

> AI does not merely automate programming. It can change the unit of software authorship from implementation production toward specification, constraint design, verification, and selection among machine-generated implementations.

This is a hypothesis-driven repository. The thesis is **not treated as established fact**. Claims must earn support through reproducible experiments and retained evidence.

## Research tracks

- [Specification primacy](research/specification-primacy.md)
- [Programming-language selection](research/language-selection.md)
- [Legacy preservation](research/legacy-preservation.md)
- [AI-assisted modding](research/ai-modding.md)
- [Tool legitimacy gap](research/tool-legitimacy-gap.md)

## Foundation

Phase 0 defines the research contract:

- [Hypotheses](HYPOTHESES.md) — six falsifiable hypotheses.
- [Terminology](TERMINOLOGY.md) — stable definitions for project terms.
- [Methodology](METHODOLOGY.md) — common experiment and evidence rules.
- [Invariants](INVARIANTS.md) — non-negotiable research constraints.
- [Roadmap](ROADMAP.md) — implementation sequence from foundation to archival record.

The repository intentionally separates **motivation**, **hypothesis**, **measurement**, **evidence**, and **conclusion**.

## Phase 0 validation

The foundational contract uses Python 3.12 or 3.14 and a pinned CommonMark parser.
Install the two hashed runtime dependencies before running the validator or tests:

```bash
python3 -m venv .venv
. .venv/bin/activate
python3 -m pip install --require-hashes -r requirements.txt
python3 scripts/validate_phase0.py
python3 -m unittest discover -s tests -v
```

The validator checks required files, exact ordered hypothesis/invariant/terminology
heading inventories, two required prose statements, the Phase 0 roadmap heading,
and a non-whitespace source-size heuristic. It does not assess the completeness
of section bodies or establish empirical support for any hypothesis.

[Validation policy](docs/VALIDATION.md) defines eligible headings, decoded prose,
excluded metadata, parser compatibility corrections, and regression coverage.

## Planned experimental flow

```text
human intent
|
v
specification
|
v
constraints / contracts
|
v
machine-generated candidate implementation
|
v
compiler + tests + static / formal checks
|
v
retained evidence
|
v
supported, weakened, or falsified claim
```

## Current status

**Phase 0 — Foundational Research Contract**

No language ranking, legacy-survival claim, modding claim, or social-perception claim is considered established merely because it appears in project motivation. Empirical phases begin after the research contract is merged.

## Scope

The project is interested in:

- spec-driven and contract-driven AI development;
- compiler and toolchain feedback as machine guidance;
- machine-verifiable programming-language properties;
- legacy-language maintenance and preservation;
- implementation fungibility across languages and toolchains;
- AI-assisted game modification and mechanic recombination;
- disclosure effects on perceptions of AI-assisted work.

The project is **not** a leaderboard for preferred programming languages and is not designed to prove that AI-generated software is inherently better or worse than human-authored software.

## License

See [LICENSE](LICENSE).
Loading
Loading