Skip to content

RFC: Audit eval harness agent loading architecture #241

Description

@evansenter

Summary

Audit how the eval harness loads agents to ensure we don't require all agents to be compiled into the eval binary.

Motivation

Current concern: If the eval binary statically links all agents, it creates:

  1. Compile-time coupling - Adding a new agent requires recompiling the eval tool
  2. Binary bloat - Eval binary includes all agent code even when testing one
  3. Slow iteration - Changes to any agent require full rebuild

The eval harness should be able to:

  • Evaluate agents without compiling them into the same binary
  • Support dynamic agent loading or external agent binaries
  • Allow third-party agents to be evaluated without source access

Scope

  1. Analyze current architecture

    • How does gemicro-eval reference agents?
    • What's the compile-time vs runtime boundary?
    • How are agents instantiated during evaluation?
  2. Evaluate alternatives

    • Agent as subprocess (JSON protocol)
    • Dynamic loading (dlopen/dylib)
    • Registry with optional agent features
    • Binary-per-agent approach
  3. Consider trade-offs

    • Complexity vs flexibility
    • Performance implications
    • Development experience
    • CI/CD impact
  4. Propose path forward

    • Minimal viable decoupling
    • Extension points for future needs

Related

  • CLI agent loading RFC (similar concerns)
  • Agent isolation principle in CLAUDE.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions