Skip to content

Add BH #2D hardware evidence harness and lightweight N-body home page - #24

Merged
EmergentMonk merged 82 commits into
mainfrom
feature/bh2d-hardware-sweep
Sep 27, 2026
Merged

EmergentMonk merged 82 commits into
mainfrom
feature/bh2d-hardware-sweep

Conversation

@EmergentMonk

@EmergentMonk EmergentMonk commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Browser N-body upgrade and latest review fixes

The Pages home page now runs real planar Barnes–Hut self-gravity with 768 interacting bodies by default, selectable from 128–2,048. Glow sprites and short position-history trails supply visual density without adding simulated mass. It includes encounter/disc/collapse presets, orbit/zoom controls, pause/single-step, a tree overlay and a pause-and-audit action. Physics timing is independent of monitor refresh rate; reduced-motion preference starts paused. The previous rotation-law instrument is preserved at rotation-lab.html, and the detailed tree lab remains at barnes-hut.html.

The eight latest review findings are addressed: isolated Python launch, sanitized system build PATH, selected toolchain hashes behind symlinked/hardlinked rustup proxies, raw-byte run logs, finite/ranged CLI validation and strict manifest JSON, normalized integer-limit parsing errors, cleared Git selectors with explicit checkout binding, and repository-relative CARGO_HOME resolution.

Validation:

  • 38 host-side harness tests and dry-run command construction passed.
  • Browser solver, UFF/Wasm, original application and new N-body integration suites passed.
  • New integration coverage confirms identical trajectories at 60 Hz and 144 Hz, pause/step, camera isolation, audit invalidation, hidden tabs and reduced motion.
  • Chromium desktop (1440 px) and mobile (390 px): all presets, controls, force audit and preserved pages checked; no page errors or horizontal mobile overflow.
  • Static Pages bundle built and local/remote Git tree hashes matched.
  • No real GPU hardware sweep was performed; hardware evidence remains pending.
  • Pages will publish after merge through the existing main-branch workflow.

Summary

Implements the remaining N-body-related deferred work that is actionable in-repo after merged PR #23: a fail-closed real-hardware BH #2D scaling/evidence harness.

The historical PE #15 deferred backlog remains untouched because it belongs to the prescribed-field/CPU-memory programme rather than the resident Barnes–Hut self-gravity line.

What this adds

  • scripts/bench-bh2d-hardware.py

    • defaults to a 512 → 65,536 resident-body sweep;
    • requires galaxy-bh-gpu-tree-parallel --require-hardware;
    • refuses dirty tracked source and pre-existing evidence directories;
    • pins the sweep to one Git revision;
    • validates every BH #2D receipt before accepting it;
    • requires hardware measurement classification and the frozen zero-host-rebuild/readback invariants;
    • checks direct-force, trajectory/oracle status, repeat-tree determinism, benchmark sample counts and BH #2C comparison/skip boundaries;
    • requires one adapter identity across the sweep;
    • SHA-256 hashes each receipt and log;
    • writes a galaxy.bh2d-hardware-scaling-manifest.v1 manifest;
    • preserves a failed manifest when a sweep point fails.
  • host-side unit tests for the sweep receipt contract;

  • native-GPU CI coverage for the tests and dry-run command construction;

  • README, roadmap and Barnes–Hut/GPU runtime documentation.

Claim boundary

This PR does not invent or claim hardware performance evidence. Mesa/software Vulkan remains validation-only.

BH #2D remains evidence-pending until the harness is actually run on a real GPU. A completed manifest is evidence only for its recorded source, workload and adapter; BH #2E production promotion remains a separate decision.

Deliberately still deferred

The historical PE #15 backlog—heterogeneous prescribed-field CPU+GPU execution, NUMA work, multi-GPU deterministic sharding, output pipelines, symmetry compression, etc.—is unchanged.

Summary by Sourcery

Establish a fail-closed harness for capturing and validating reproducible real-GPU BH #2D scaling evidence without making automatic performance or production-promotion claims.

New Features:

  • Add a fail-closed runner for collecting source-pinned BH #2D real-hardware scaling evidence across configurable resident-body sizes.
  • Generate manifests that record workload, adapter, toolchain, source revision, and SHA-256 hashes for receipts and logs, including failed sweeps.

Enhancements:

  • Validate hardware classification, correctness and oracle gates, deterministic tree rebuilding, benchmark samples, allocation bounds, and BH #2C comparison boundaries for every sweep point.
  • Enforce clean source and build environments, stable toolchain and Cargo configuration, fresh output directories, and a single adapter identity throughout each sweep.

CI:

  • Run BH #2D sweep contract tests and dry-run command validation in native-GPU CI.

Documentation:

  • Document the real-hardware BH #2D evidence workflow, its claim boundaries, and evidence-pending status in the README, roadmap, and GPU runtime documentation.

Tests:

  • Add comprehensive host-side tests for particle validation, provenance checks, receipt contracts, oracle skip rules, determinism, benchmark validation, and failure conditions.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @EmergentMonk, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 1 day and 21 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Sep 27, 2026

Copy link
Copy Markdown

Reviewer's Guide

This PR introduces a fail-closed BH #2D real-hardware scaling harness that executes the existing GPU verifier across a source-pinned resident-body sweep, enforces correctness, determinism, hardware, oracle, and benchmark invariants, and emits hashed complete or failed manifests; accompanying tests, CI coverage, and documentation make the evidence workflow reproducible without claiming performance or production promotion.

Sequence diagram for the BH #2D hardware evidence sweep

sequenceDiagram
    participant Runner as bench-bh2d-hardware.py
    participant Git as Git
    participant Verifier as galaxy-bh-gpu-tree-parallel
    participant GPU as Real GPU
    participant Manifest as manifest.json

    Runner->>Runner: parse_particles()
    Runner->>Git: git diff --quiet
    Runner->>Git: git diff --cached --quiet
    Runner->>Git: git rev-parse HEAD
    loop Each resident-body count
        Runner->>Verifier: run with --require-hardware
        Verifier->>GPU: execute BH #2D workload
        GPU-->>Verifier: receipt.json and runtime output
        Verifier-->>Runner: completed receipt
        Runner->>Runner: validate_receipt()
        Runner->>Runner: sha256_file(receipt)
        Runner->>Runner: sha256_file(log)
        Runner->>Manifest: append validated run and hashes
    end
    Runner->>Manifest: mark status complete
Loading

Flow diagram for fail-closed BH #2D evidence validation

flowchart TD
    A[Start sweep] --> B[Validate particle counts and options]
    B --> C{Dry run?}
    C -->|Yes| D[Print exact verifier commands]
    C -->|No| E[Require clean tracked source]
    E --> F[Pin Git revision and create manifest]
    F --> G[Run verifier with --require-hardware]
    G --> H{Receipt passes hardware and correctness gates?}
    H -->|No| I[Write failed manifest and stop]
    H -->|Yes| J{Adapter matches prior runs?}
    J -->|No| I
    J -->|Yes| K[Hash receipt and log]
    K --> L{More sweep points?}
    L -->|Yes| G
    L -->|No| M[Write complete source-pinned manifest]
Loading

File-Level Changes

Change Details Files
Adds a fail-closed, source-pinned real-GPU BH #2D scaling sweep runner that validates receipts and preserves auditable evidence.
  • Runs a strictly increasing resident-body sweep through the existing GPU verifier with hardware enforcement.
  • Rejects dirty tracked trees and pre-existing output directories, and pins all runs to one Git revision.
  • Validates receipt schema, hardware classification, correctness thresholds, zero host rebuild/readback invariants, determinism, oracle status, benchmark samples, and BH #2C comparison boundaries.
  • Requires a stable adapter identity and records SHA-256 hashes for every receipt and log.
  • Writes incremental complete or failed manifests with workload, host, source, run summaries, and claim-boundary metadata.
scripts/bench-bh2d-hardware.py
Adds host-side contract tests covering sweep parsing and the principal receipt acceptance and rejection gates.
  • Tests strict particle-list validation.
  • Accepts valid hardware receipts at and above the oracle limit with the expected BH #2C behavior.
  • Rejects software measurements, repeat-tree mismatches, and invalid BH #2C skip boundaries.
tests/test_bh2d_hardware_sweep.py
Integrates the harness into native-GPU CI and documents its evidence-only role.
  • Runs receipt-contract tests and dry-run command construction in native-GPU CI.
  • Documents invocation, default sweep sizes, hardware-only requirements, manifest semantics, and non-promotion claim boundaries.
  • Updates roadmap and runtime guidance while keeping BH #2D evidence-pending until a real-GPU run exists.
.github/workflows/native-gpu.yml
README.md
ROADMAP.md
docs/BARNES-HUT.md
docs/GPU-RUNTIME.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T16:47:36.513908Z da9bf28 Manual request
🔒 Security Review ✅ Completed 2026-09-27T06:29:43.529888Z d142ca8 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copy link
Copy Markdown
Member Author

@codex please review and security review the exact current head d142ca8. Focus on actionable correctness or security defects in the BH #2D hardware sweep harness, receipt validation, evidence provenance, path handling, and CI integration. Report executed vs statically inferred reproductions where applicable.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d142ca835a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py Outdated

Copy link
Copy Markdown
Member Author

@codex please review and security review the exact current head 2570e34. The seven earlier P2 findings are addressed with regression coverage and their threads are resolved. Report only new actionable correctness/security defects on this exact SHA; do not repeat fixed findings without a new failing case.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2570e34baa

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py Outdated

Copy link
Copy Markdown
Member Author

@codex please review and security review the exact current head a8d444c. The latest five P2 findings are fixed with dedicated regressions and their threads are resolved. Report only new actionable correctness/security defects on this exact SHA; do not repeat fixed findings without a new failing case.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a8d444cf32

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py Outdated
Comment thread scripts/bench-bh2d-hardware.py

Copy link
Copy Markdown
Member Author

@codex please review and security review the exact current head d00d223. The latest four P2 findings are fixed with dedicated regressions and their threads are resolved. Report only new actionable correctness/security defects on this exact SHA; do not repeat fixed findings without a new failing case.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d00d223e15

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py Outdated

Copy link
Copy Markdown
Member Author

@codex please review and security review the exact current head e9f0145. The latest two P1 and two P2 findings are fixed with dedicated regressions and their threads are resolved. Report only new actionable correctness/security defects on this exact SHA; do not repeat fixed findings without a new failing case.

@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cb46356587

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware.py
Comment thread scripts/bench-bh2d-hardware.py Outdated
@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 682369c4dd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/bench-bh2d-hardware-launch.cmd Outdated
Comment thread docs/GPU-RUNTIME.md Outdated
Comment thread scripts/bench-bh2d-hardware.py
@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex Please Review this exact SHA da9bf28 against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: da9bf28d6c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@EmergentMonk
EmergentMonk merged commit e051324 into main Sep 27, 2026
36 checks passed
@EmergentMonk
EmergentMonk deleted the feature/bh2d-hardware-sweep branch September 27, 2026 16:33

Copy link
Copy Markdown
Member Author

@codex Please Review this exact SHA e051324 against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e051324b12

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +386 to +393
def build_environment(context: dict[str, Any]) -> dict[str, str]:
env = {
key: value
for key, value in os.environ.items()
if not key.startswith("GIT_")
and key not in BUILD_ENV_EXACT
and not any(pattern.fullmatch(key) for pattern in BUILD_ENV_PATTERNS)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bind inherited MSVC build environment inputs

Minimal reproduction (executed at helper level on this commit): set LIB=C:\attacker-libs, INCLUDE=C:\attacker-includes, LIBPATH=C:\attacker-netmodules, CL=/DCHANGED_BUILD, and _CL_=/link /DEBUG; build_environment_overrides() returns [], while build_environment() preserves all five values. Expected: native Windows capture rejects these inputs or records and binds them because MSVC uses them for library/header lookup and injected compiler/linker options. Actual: the launcher inherits them, these lines pass them to Cargo, and the manifest still reports environment_overrides: [], so a Windows verifier can be built from unrecorded inputs. The resulting native build impact is statically inferred because the reproduction host is Linux.

Useful? React with 👍 / 👎.

Comment on lines +1137 to +1139
require(
leaf_count <= particles,
"receipt.tree.leaf_count cannot exceed the resident particle count",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require fewer leaves than particles after a bucket split

Fresh executed reproduction beyond the previously fixed topology ceilings: the checked-in 512-particle fixture with active_cell_count=1023, leaf_count=512, and max_depth=9 still returns a completed summary. Expected: rejection, because with a four-particle bucket any deepest internal cell contains at least five particles but has at most four leaf children, forcing at least one leaf to contain multiple particles and therefore leaf_count < particles. Actual: this non-strict bound accepts an impossible all-singleton-leaf tree as hardware evidence.

Useful? React with 👍 / 👎.

Comment thread runtime/src/nbody_gpu.rs
Comment on lines +302 to 305
"adapter_count": adapter_count,
"name": info.name,
"backend": format!("{:?}", info.backend),
"device_type": format!("{:?}", info.device_type),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Include hardware IDs in adapter identity

Executed identity-level reproduction: two GPU objects with identical index, count, name, backend, type, and driver strings but different vendor and device IDs produce exactly the same adapter_identity(). Expected: they are distinct adapters because the completed manifest claims one fixed adapter for the sweep. Actual: wgpu::AdapterInfo's hardware IDs are omitted here and from the validator identity, so if separate sweep processes enumerate different cards under the same generic name/driver strings and index, the adapter-change check passes; that multi-adapter impact is statically inferred.

Useful? React with 👍 / 👎.

Comment on lines +608 to +612
executable = bool(candidate.stat().st_mode & stat.S_IXUSR)
if executable != (mode == "100755"):
raise SweepError(
f"tracked source tree is dirty; executable mode changed: {relative}"
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Skip POSIX executable-bit checks on native Windows

Statically inferred native-Windows reproduction, with the repository precondition executed: git ls-files --stage reports multiple tracked .sh and .py files as mode 100755, including scripts/bench-cpu-memory-wall.sh. Expected: a clean Windows checkout accepted by Git can start the documented hardware sweep. Actual: Windows filesystems do not preserve that POSIX executable bit for these script extensions, so candidate.stat().st_mode & S_IXUSR is false and this comparison raises executable mode changed before any run begins. Make this mode check platform-aware while retaining the raw-content comparison.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant