From 7c72202f1358be0b376f9ca40e8104a1871f8fc8 Mon Sep 17 00:00:00 2001 From: spiffytech Date: Thu, 12 Mar 2026 10:13:22 -0400 Subject: [PATCH] Added Ollama Cloud benchmarks --- README.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/README.md b/README.md index cdeb4ce..af9e320 100644 --- a/README.md +++ b/README.md @@ -43,6 +43,11 @@ add more provider results! |Parasail |GLM-4.7 |:x: 83%| |Parasail |Kimi K2 Thinking|:x: 75%| +|Provider |Model |Success Rate| +|---------|----------------|------------| +|Ollama Cloud |GLM-4.7 |:x: 88%| +|Ollama Cloud |Minimax M2|:x: 62%| + Note for attempting reproductions: generally all tests are reproducible with `--count 1` and `--count 1 --stream`, but for evaluating the response-in-reasoning eval, you generally will need a high count to reproduce