Summary
Audit how the eval harness loads agents to ensure we don't require all agents to be compiled into the eval binary.
Motivation
Current concern: If the eval binary statically links all agents, it creates:
- Compile-time coupling - Adding a new agent requires recompiling the eval tool
- Binary bloat - Eval binary includes all agent code even when testing one
- Slow iteration - Changes to any agent require full rebuild
The eval harness should be able to:
- Evaluate agents without compiling them into the same binary
- Support dynamic agent loading or external agent binaries
- Allow third-party agents to be evaluated without source access
Scope
-
Analyze current architecture
- How does
gemicro-eval reference agents?
- What's the compile-time vs runtime boundary?
- How are agents instantiated during evaluation?
-
Evaluate alternatives
- Agent as subprocess (JSON protocol)
- Dynamic loading (dlopen/dylib)
- Registry with optional agent features
- Binary-per-agent approach
-
Consider trade-offs
- Complexity vs flexibility
- Performance implications
- Development experience
- CI/CD impact
-
Propose path forward
- Minimal viable decoupling
- Extension points for future needs
Related
- CLI agent loading RFC (similar concerns)
- Agent isolation principle in CLAUDE.md
Summary
Audit how the eval harness loads agents to ensure we don't require all agents to be compiled into the eval binary.
Motivation
Current concern: If the eval binary statically links all agents, it creates:
The eval harness should be able to:
Scope
Analyze current architecture
gemicro-evalreference agents?Evaluate alternatives
Consider trade-offs
Propose path forward
Related