Skip to content

Public leaderboard + publish fleet-model scored results #1

Description

@ThinkOffApp

From claudemm's repo review (2026-07-06). OMATS is our most differentiated but under-marketed asset (★2) — the 5-stage capability framework tests multi-agent failure modes that MMLU/SWE-bench don't.

Goal: turn OMATS from a spec into a public, citable benchmark.

Scope:

  • Run the current fleet models (gpt-5.5, kimi-k2.5, glm-5.2, gemini-3.5-flash, claude, etc.) through the scripted OMATS room scenarios and produce graded scorecards.
  • Build a public leaderboard page (static site is fine) that renders the scorecards per stage (1–5).
  • Write it up as a launch post / short paper referencing the 5-stage framework.
  • Link the leaderboard from the repo README + thinkoff.io.

Biggest marketing ROI item in the review — it showcases the whole multi-agent stack.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions