Skip to content

fix: resolve LLM judge overall_score overriding weighted benchmark score - #8

Closed
RealDiligent wants to merge 1 commit into
gittensor-agent-forge:mainfrom
RealDiligent:fix/critical-issue-judge-score-override
Closed

fix: resolve LLM judge overall_score overriding weighted benchmark score#8
RealDiligent wants to merge 1 commit into
gittensor-agent-forge:mainfrom
RealDiligent:fix/critical-issue-judge-score-override

Conversation

@RealDiligent

Copy link
Copy Markdown

Summary

  • Stop evaluate_openrouter_vision() from replacing the weighted dimension aggregate with the LLM self-reported overall_score field.
  • Add regression test ensuring inflated self-reported overall scores are ignored.

Test plan

  • python -m pytest tests/test_scoring.py -v

@AceronX AceronX closed this Jul 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants