Skip to content

[codex] add AI participation foundation evaluator - #6

Merged
s133per merged 1 commit into
mainfrom
codex/add-ai-participation-foundation-evaluator
Jun 14, 2026
Merged

[codex] add AI participation foundation evaluator#6
s133per merged 1 commit into
mainfrom
codex/add-ai-participation-foundation-evaluator

Conversation

@s133per

@s133per s133per commented Jun 14, 2026

Copy link
Copy Markdown
Contributor

What Changed

  • Added a new foundation-level ai-native-ai-participation-evaluator weighted at 40% of the foundation score.
  • Added six child evaluators for agent thread participation, source-control AI participation, skill activation depth, AI self-assessment loops, human follow-through, and human-AI collaboration traces.
  • Updated issue and PR templates so work must record AI trigger quality, skill usage, human follow-through, skipped workflow steps, and agent pushback.
  • Regenerated the foundation self-evaluation report and docs baseline to include the new AI participation dimension.
  • Expanded packaging tests to enforce the new evaluator tree, weight ratio, template fields, and recent-change rubric coverage.

Intentionally Not Done

  • Did not run live agent evals. This change updates evaluator contracts, templates, and persisted self-evaluation artifacts rather than validating a real agent runtime behavior change.
  • Did not alter product-domain code or add hosted/runtime service behavior.

Command Evidence

  • node --test tests/*.test.mjs
  • pnpm self-eval:validate
  • pnpm --dir .agents/skills/ai-native-eval/scripts/eval test:tool
  • pnpm self-eval:render
  • pnpm self-eval:check
  • pnpm test
  • pnpm skill-eval:contract
  • git diff --check

Skill Contracts Or Rubrics Changed

  • Added a new foundation group evaluator and six leaf evaluator contracts.
  • Updated the foundation evaluator manifest so AI participation contributes 40% of the foundation score.
  • Each new leaf evaluator scores configuration evidence and recent issue/PR/thread execution evidence, with unavailable GitHub/source-control access treated as absent evidence.

Self-Evaluation Artifact Changed

  • Updated self-evaluations/foundation-20260614/report.md.
  • Added six new evaluator output JSON files under self-evaluations/foundation-20260614/run/evaluators/.
  • Updated README/docs baseline score to 2.8 / 10, level 2, reflecting low current execution evidence for the new AI participation rules.

Human-Only Review Decisions

  • Reviewers should confirm that 40% is the intended foundation weight and that the issue/PR template prompts are strict enough for future process evidence.

Skipped Gates

  • Live agent eval was intentionally deferred because no real-agent behavior path changed in this PR.

@s133per
s133per marked this pull request as ready for review June 14, 2026 18:53
@s133per
s133per merged commit e1d5173 into main Jun 14, 2026
1 check passed
@s133per
s133per deleted the codex/add-ai-participation-foundation-evaluator branch June 14, 2026 18:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant