Skip to content

Commit 0915691

Browse files
sophiasophia
authored andcommitted
Sync GLM model pricing docs
1 parent c1b93db commit 0915691

12 files changed

Lines changed: 12 additions & 113 deletions

File tree

docs/llmservice/models/glm-5-1.md

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22

33
## Overview
44

5-
GLM-5.1 is an open-source flagship AI model developed by Z.ai, formerly Zhipu AI, a Tsinghua University spinoff and the first publicly traded foundation model company. Released on April 7, 2026, it is a post-training upgrade to GLM-5, built on a 754-billion-parameter Mixture-of-Experts architecture with 40 billion active parameters per token. GLM-5.1 is designed for agentic engineering and long-horizon autonomous software development.
5+
GLM-5.1 is an open-source flagship AI model developed by Z.ai, formerly Zhipu AI, a Tsinghua University spinoff and the first publicly traded foundation model company. Released on April 7, 2026, it is a post-training upgrade in the GLM family, built on a 754-billion-parameter Mixture-of-Experts architecture with 40 billion active parameters per token. GLM-5.1 is designed for agentic engineering and long-horizon autonomous software development.
66

77
## Key Features
88

@@ -23,7 +23,7 @@ GLM-5.1 is an open-source flagship AI model developed by Z.ai, formerly Zhipu AI
2323
| :----------------- | :----------------------------------------------------------------------------------------------------------- |
2424
| **Reasoning** | AIME 2026: 95.3%, GPQA-Diamond: 86.2%, with strong system-level reasoning across planning and iterative debugging |
2525
| **Coding** | SWE-Bench Pro 58.4%, CyberGym 68.7%, BrowseComp 68.0%, MCP-Atlas 71.8% |
26-
| **Multimodal** | Text only. No image, audio, or video input. A separate GLM-5V-Turbo variant is available for vision tasks |
26+
| **Multimodal** | Text only. No image, audio, or video input. Vision tasks require a separate vision-capable model |
2727
| **Response Speed** | Not independently benchmarked yet; expected to be comparable to similar-scale MoE models |
2828
| **Context Window** | 200K tokens |
2929
| **Max Output** | 128K tokens |
@@ -32,7 +32,7 @@ GLM-5.1 is an open-source flagship AI model developed by Z.ai, formerly Zhipu AI
3232

3333
### Known Limitations
3434

35-
* Text-only input with no native multimodal support. Vision use cases rely on the separate GLM-5V-Turbo model.
35+
* Text-only input with no native multimodal support. Vision use cases rely on a separate vision-capable model.
3636
* Math and science benchmark scores trail some top proprietary models, making it less suitable for purely quantitative research tasks.
3737
* On broader coding composites such as Terminal-Bench 2.0 plus NL2Repo, Claude Opus 4.6 still leads.
3838
* Self-hosting requires substantial compute resources because of the 754B parameter count.

docs/llmservice/models/glm-5-2.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -24,7 +24,7 @@ GLM-5.2 is a GLM-family text foundation model developed by Z.AI and released on
2424
| **Reasoning** | Supports deep-thinking mode and `reasoning_effort`; Z.AI positions it for complex engineering, debugging, and long-chain reasoning workflows. |
2525
| **Creative Writing** | Supports general text generation through the chat completion API, but official GLM-5.2 materials emphasize coding and engineering use cases. |
2626
| **Coding** | Z.AI reports Terminal-Bench 2.1 score of 81.0 and SWE-bench Pro score of 62.1, with focus on long-horizon coding-agent scenarios. |
27-
| **Multimodal** | Text input and text output. Vision and multimodal workflows are handled by separate Z.AI models such as GLM-5V-Turbo. |
27+
| **Multimodal** | Text input and text output. Vision and multimodal workflows are handled by separate Z.AI vision-language models. |
2828
| **Response Speed** | Official docs do not publish latency or tokens-per-second figures; streaming responses and streaming tool calls are supported. |
2929
| **Context Window** | 1M tokens. |
3030
| **Max Output** | 128K tokens. |
@@ -33,7 +33,7 @@ GLM-5.2 is a GLM-family text foundation model developed by Z.AI and released on
3333

3434
### Known Limitations
3535

36-
* Text-only model; image, video, and GUI-understanding tasks require a separate vision-language model such as GLM-5V-Turbo.
36+
* Text-only model; image, video, and GUI-understanding tasks require a separate vision-language model.
3737
* Very long contexts and 128K outputs can increase latency and cost; cap `max_tokens` and use context caching where applicable.
3838

3939
## Credits Usage

docs/llmservice/models/glm-5.md

Lines changed: 0 additions & 47 deletions
This file was deleted.

docs/llmservice/models/minimax-m2.7.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ MiniMax M2.7 is a reasoning-focused large language model developed by MiniMax (S
66

77
## Key Features
88

9-
* **Cost-Efficient Reasoning**: Achieves intelligence scores comparable to GLM-5 and Kimi K2.5 while costing roughly one-third as much to run, with 20% fewer output tokens needed for equivalent results.
9+
* **Cost-Efficient Reasoning**: Delivers strong reasoning performance at a lower operating cost than many flagship models, with 20% fewer output tokens needed for equivalent results.
1010
* **Low Hallucination Rate**: Scores a 34% hallucination rate on the AA-Omniscience Index, lower than Claude Sonnet 4.6 and Gemini 3.1 Pro Preview.
1111
* **Multi-Agent Collaboration**: Native support for multi-agent orchestration and complex skill coordination, including dynamic tool discovery and invocation at runtime.
1212
* **Self-Evolution**: Can autonomously complete 30-50% of reinforcement-learning research workflows, representing an early step toward model self-improvement.

docs/llmservice/pricing-and-usage.md

Lines changed: 0 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,6 @@ The platform uses a unified Credits system to measure and settle usage across al
2020
| Qwen3.6-27B | 0.19 | 0.19 | 0.019 | 2.99 | - |
2121
| GLM-5.2 | 1.40 | 1.40 | 0.28 | 4.40 | - |
2222
| GLM-5.1 | 1.40 | 1.40 | 0.26 | 4.40 | - |
23-
| GLM-5 | 1.00 | 1.00 | 0.20 | 3.20 | - |
2423
| DeepSeek V3.2 | 0.29 | 0.29 | 0.145 | 0.44 | - |
2524
| DeepSeek V4 Flash | 0.28 | 0.28 | 0.0056 | 0.56 | - |
2625
| DeepSeek V4 Pro | 0.87 | 0.87 | 0.0087 | 1.74 | - |

i18n/zh-Hans/docusaurus-plugin-content-docs/current/llmservice/models/glm-5-1.md

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22

33
## 概述
44

5-
GLM-5.1 是由 Z.ai 开发的开源旗舰 AI 模型。Z.ai 前身为智谱 AI,源自清华大学,也是首家公开上市的基础模型公司。GLM-5.1 于 2026 年 4 月 7 日发布,是在 GLM-5 基础上的一次后训练升级,采用 7540 亿参数的 Mixture-of-Experts 架构,每个 token 激活约 400 亿参数,重点面向 Agent 工程与长周期自主软件开发场景。
5+
GLM-5.1 是由 Z.ai 开发的开源旗舰 AI 模型。Z.ai 前身为智谱 AI,源自清华大学,也是首家公开上市的基础模型公司。GLM-5.1 于 2026 年 4 月 7 日发布, GLM 系列的一次后训练升级,采用 7540 亿参数的 Mixture-of-Experts 架构,每个 token 激活约 400 亿参数,重点面向 Agent 工程与长周期自主软件开发场景。
66

77
## 核心特性
88

@@ -23,7 +23,7 @@ GLM-5.1 是由 Z.ai 开发的开源旗舰 AI 模型。Z.ai 前身为智谱 AI,
2323
| :--- | :--- |
2424
| **推理能力** | AIME 2026:95.3%,GPQA-Diamond:86.2%,在规划与迭代调试场景中具备较强的系统级推理能力 |
2525
| **编程能力** | SWE-Bench Pro 58.4%,CyberGym 68.7%,BrowseComp 68.0%,MCP-Atlas 71.8% |
26-
| **多模态能力** | 仅支持文本,不支持图像、音频或视频输入;视觉场景可使用单独的 GLM-5V-Turbo 变体 |
26+
| **多模态能力** | 仅支持文本,不支持图像、音频或视频输入;视觉场景需要使用单独的视觉能力模型 |
2727
| **响应速度** | 暂无独立公开测速结果,预计与同规模 MoE 模型相近 |
2828
| **上下文窗口** | 200K tokens |
2929
| **最大输出** | 128K tokens |
@@ -32,7 +32,7 @@ GLM-5.1 是由 Z.ai 开发的开源旗舰 AI 模型。Z.ai 前身为智谱 AI,
3232

3333
### 已知限制
3434

35-
* 仅支持文本输入,不具备原生多模态能力;视觉任务需依赖独立的 GLM-5V-Turbo 模型
35+
* 仅支持文本输入,不具备原生多模态能力;视觉任务需依赖独立的视觉能力模型
3636
* 数学与科学基准成绩仍落后于部分顶级专有模型,因此在纯量化研究任务上不一定是最优选择。
3737
* 在更广泛的编码综合评测(如 Terminal-Bench 2.0 + NL2Repo)中,Claude Opus 4.6 仍然领先。
3838
* 由于参数规模达到 754B,自托管需要较高的计算资源。

i18n/zh-Hans/docusaurus-plugin-content-docs/current/llmservice/models/glm-5-2.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -24,7 +24,7 @@ GLM-5.2 是由 Z.AI 开发的 GLM 系列文本基础模型,于 2026 年 6 月
2424
| **推理能力** | 支持 deep-thinking 模式和 `reasoning_effort`;Z.AI 将其定位于复杂工程、调试和长链路推理工作流 |
2525
| **创意写作** | 支持通过 chat completion API 进行通用文本生成,但官方 GLM-5.2 材料更强调代码和工程场景 |
2626
| **编程能力** | Z.AI 报告 Terminal-Bench 2.1 得分为 81.0,SWE-bench Pro 得分为 62.1,重点面向长周期 Coding Agent 场景 |
27-
| **多模态能力** | 文本输入和文本输出;视觉和多模态工作流由 GLM-5V-Turbo 等独立 Z.AI 模型处理 |
27+
| **多模态能力** | 文本输入和文本输出;视觉和多模态工作流由独立的 Z.AI 视觉语言模型处理 |
2828
| **响应速度** | 官方文档未公布延迟或 tokens-per-second 数据;支持流式响应和流式工具调用 |
2929
| **上下文窗口** | 1M tokens |
3030
| **最大输出** | 128K tokens |
@@ -33,7 +33,7 @@ GLM-5.2 是由 Z.AI 开发的 GLM 系列文本基础模型,于 2026 年 6 月
3333

3434
### 已知限制
3535

36-
* 该模型为文本模型;图像、视频和 GUI 理解任务需要使用 GLM-5V-Turbo 等独立视觉语言模型
36+
* 该模型为文本模型;图像、视频和 GUI 理解任务需要使用独立视觉语言模型
3737
* 超长上下文和 128K 输出可能增加延迟和成本;建议按需限制 `max_tokens`,并在适用场景使用上下文缓存。
3838

3939
## 积分消耗

i18n/zh-Hans/docusaurus-plugin-content-docs/current/llmservice/models/glm-5.md

Lines changed: 0 additions & 50 deletions
This file was deleted.

i18n/zh-Hans/docusaurus-plugin-content-docs/current/llmservice/models/minimax-m2.7.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ MiniMax M2.7 是由 MiniMax(上海)开发的推理型大语言模型,于 2
66

77
## 核心特性
88

9-
* **高性价比推理能力**在智能表现上可与 GLM-5 和 Kimi K2.5 对标,但运行成本约为其三分之一,同时在等效任务下所需输出 token 减少约 20%。
9+
* **高性价比推理能力**在较低运行成本下提供较强推理表现,同时在等效任务下所需输出 token 减少约 20%。
1010
* **较低幻觉率**:在 AA-Omniscience Index 上的幻觉率为 34%,低于 Claude Sonnet 4.6 和 Gemini 3.1 Pro Preview。
1111
* **多 Agent 协作**:原生支持多 Agent 编排和复杂技能协同,包括运行时动态工具发现与调用。
1212
* **自进化能力**:可自主完成约 30%-50% 的强化学习研究工作流,体现出早期的模型自我改进能力。

i18n/zh-Hans/docusaurus-plugin-content-docs/current/llmservice/pricing-and-usage.md

Lines changed: 0 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,6 @@
2020
| Qwen3.6-27B | 0.19 | 0.19 | 0.019 | 2.99 | - |
2121
| GLM-5.2 | 1.40 | 1.40 | 0.28 | 4.40 | - |
2222
| GLM-5.1 | 1.40 | 1.40 | 0.26 | 4.40 | - |
23-
| GLM-5 | 1.00 | 1.00 | 0.20 | 3.20 | - |
2423
| DeepSeek V3.2 | 0.29 | 0.29 | 0.145 | 0.44 | - |
2524
| DeepSeek V4 Flash | 0.28 | 0.28 | 0.0056 | 0.56 | - |
2625
| DeepSeek V4 Pro | 0.87 | 0.87 | 0.0087 | 1.74 | - |

0 commit comments

Comments
 (0)