A source-grounded, duplication-checked component atlas for LLM benchmark charts, model launch comparisons, independent leaderboards, and research-paper figures.
Live catalog · Product story · Static component index · X / @logiclogic1223 · Chinese README · Component contract · Browser API · JSON Schema · Source inventory
Benchmark screenshots are abundant, but reusable chart grammar is not. The same bar chart is frequently copied into several folders, recolored for a new vendor, and counted again. Benchmark Atlas uses a stricter unit of value:
- 83 live SVG components generated from structured data
- 83 unique chart grammars, renderers, and visual-system identifiers
- 42 source lineages across vendors, leaderboards, labs, and arXiv papers
- 10 chart families from ranking and uncertainty to agent trajectories
- search, multi-axis filtering, enlarged inspection, JSON access, and SVG export
- validation that fails on duplicates, missing sources, invalid SVG values, or missing accessibility metadata
Changing only a color, model name, or file format does not create a new component.
No build step or runtime dependency is required.
git clone https://github.com/AidenNovak/llm-benchmark-atlas.git
cd llm-benchmark-atlas
npm run validate
npm run serveOpen http://127.0.0.1:4173/library/.
The public site is normally deployed from the gh-pages branch. Maintainers can
publish the validated static bundle without a hosted runner:
npm run deploy:pagesRun the same interaction suite against production with:
npm run qa -- https://aidennovak.github.io/llm-benchmark-atlas/| Family | Representative components |
|---|---|
| Ranking and comparison | master table, lollipop rank, grouped hatch, slopegraph, bump chart |
| Scale, cost, and efficiency | compute frontier, thinking saturation, log-log scaling, Pareto, bubbles |
| Multi-dimensional profile | radar, facets, parallel coordinates, polar rose, glyph matrix |
| Distribution and uncertainty | forest interval, violin, box plot, ridgeline, ECDF, calibration |
| Diagnostics and matrices | heatmap, confusion matrix, context decay, ablation waterfall, win matrix |
| Agent and process evaluation | long-horizon ledger, token area, swimlane, Sankey, survival, solve curve |
| Special encoding and coverage | cylinders, waffle, treemap, beeswarm |
| Vendor release reproductions | Gemini triptych, OpenAI thinking pairs, Claude intervals, TML frontier |
| Figure-verified Asian labs | Kimi, MiniMax, GLM, InternLM2, ERNIE 5.0, Step-3, Yi, Hunyuan, Seed, and DeepSeek research grammars |
The source lineages include OpenAI/ChatGPT, Anthropic/Claude, Google/Gemini, DeepSeek, Qwen, Meta/Llama, Mistral, xAI/Grok, Microsoft/Phi, Amazon Nova, Cohere, NVIDIA/Nemotron, Thinking Machines Lab, LMArena, Artificial Analysis, OpenCompass, HELM, SWE-bench, Terminal-Bench, LiveCodeBench, and others.
The project deliberately stays framework-free so every chart is inspectable and exportable.
library/catalog.js source registry, component metadata, demo data
library/renderers.js 40 core pure-SVG renderers
library/vendor-series.js vendor-specific extension registry and renderers
library/research-series.js vendor detail and paper-figure extension series
library/asian-series.js figure-verified Kimi research series
library/lab-series.js figure-verified MiniMax and GLM research series
library/lab-systems-series.js InternLM2, ERNIE, Step-3, and Yi systems series
library/frontier-systems-series.js Hunyuan, Seed, and DeepSeek systems series
library/api.js stable query, render, and extension API
library/catalog.generated.json machine-readable registry snapshot
library/app.js search, filters, details, JSON copy, SVG download
scripts/validate-* contract and runtime validation
research/ source evidence and chart taxonomy
See Architecture for the data flow and extension model.
Every contribution must add material information value:
- cite a first-party, leaderboard, or paper source;
- explain the benchmark story the chart is good at telling;
- provide a new grammar, visual-system ID, and pure renderer;
- mark illustrative data clearly or add a versioned data citation;
- pass
npm run validateand mobile visual QA.
Read CONTRIBUTING.md and use the new-chart issue template before implementing a large series.
Preview values are illustrative and visibly marked DEMO DATA. Source links
document lineage; they do not imply endorsement or claim that the preview is a
current leaderboard. See NOTICE.md and
Data policy.
The next releases focus on versioned real-data adapters, an ESM package API, more verified model-lab sources, visual regression snapshots, and additional chart grammars only where a real evaluation use case justifies them. See ROADMAP.md.
The project code is available under the MIT License. Third-party names and benchmark marks remain the property of their owners.
