Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ jobs:

- uses: actions/setup-python@v5
with:
python-version: "3.13" # 对齐 pve02 线上版本
python-version: "3.13" # 对齐线上生产版本

- name: 后端单元测试
run: |
Expand Down
18 changes: 9 additions & 9 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,22 +166,22 @@ cat logs/fancontroller.log.YYYY-MM-DD

```yaml
prometheus:
base_url: "http://192.168.6.31:30091"
base_url: "http://192.0.2.20:30091"
timeout: 10

servers:
- type: epycd8
prometheus: # per-server 覆盖(多机场景必需)
fan_instance: "192.168.6.7:9290" # 风扇转速 ← ipmi_exporter
temp_instance: "192.168.6.7:9100" # CPU 温度 ← node_exporter
gpu_instance: "192.168.6.7:9400" # GPU 温度 ← DCGM exporter
fan_instance: "192.0.2.10:9290" # 风扇转速 ← ipmi_exporter
temp_instance: "192.0.2.10:9100" # CPU 温度 ← node_exporter
gpu_instance: "192.0.2.10:9400" # GPU 温度 ← DCGM exporter
```

| 数据 | 指标 | instance(pve02 实测) |
| 数据 | 指标 | instance(实测示例) |
|------|------|----------------------|
| 风扇转速 | `ipmi_fan_speed_rpm{name="FRNT_FAN1"}` | `192.168.6.7:9290` |
| CPU 温度 | `node_hwmon_temp_celsius`(按语义标签 `label="Tctl"` 过滤) | `192.168.6.7:9100` |
| GPU 温度 | `DCGM_FI_DEV_GPU_TEMP` | `192.168.6.7:9400` |
| 风扇转速 | `ipmi_fan_speed_rpm{name="FRNT_FAN1"}` | `192.0.2.10:9290` |
| CPU 温度 | `node_hwmon_temp_celsius`(按语义标签 `label="Tctl"` 过滤) | `192.0.2.10:9100` |
| GPU 温度 | `DCGM_FI_DEV_GPU_TEMP` | `192.0.2.10:9400` |

⚠️ **三个 instance 对应三个不同的 exporter,混用会直接查不到数据。** 配置项
分别为 `fan_instance` / `temp_instance` / `gpu_instance`(笼统的 `instance` 仍
Expand Down Expand Up @@ -219,7 +219,7 @@ servers:

本仓库真正在用的是 `app/`(FastAPI 控制台)+ `frontend/`(Vue3 前端):

- 控速依据是 **GPU 温度**(DCGM),目标机 pve02 的机箱风扇,写 `ipmitool raw 0x3a 0x01`
- 控速依据是 **GPU 温度**(DCGM),目标机的机箱风扇,写 `ipmitool raw 0x3a 0x01`
- 运行时状态(模式 / 管控 GPU / 分配关系 / 审计)**全部落 SQLite**,配置文件只是首次运行的种子

**实时数据源:三个外部 exporter(自备组件,不随仓库提供,装法不限)**
Expand Down
47 changes: 22 additions & 25 deletions DEPLOY.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# 部署手册 —— GPU 风扇控制台

> 目标机:**pve02**(ASRock Rack EPYCD8,192.168.6.7,BMC 2.20)
> 目标机:你的 GPU 宿主机(下文 `<目标机>` / `<连接名>` 按实际替换)
> 部署目录:`/opt/gpu-fan-console` | 服务名:`gpu-fan-console.service`
> 访问地址:`http://192.168.6.7:8765`
> 最后核对:2026-09-28(当时线上 Python 3.13.5 / fastapi 0.141.1 / uvicorn 0.53.0)
> 访问地址:`http://<目标机>:8765`
> 参考环境:Python 3.13.5 / fastapi 0.141.1 / uvicorn 0.53.0(2026-09-28 核对)

---

Expand Down Expand Up @@ -37,11 +37,11 @@ FastAPI 挂 `StaticFiles` 一起发出去。所以:
| 代码目录 | `/opt/gpu-fan-console` | 工作目录,service 的 `WorkingDirectory` |
| 虚拟环境 | `/opt/gpu-fan-console/.venv` | Python 3.13.5,已装 fastapi / uvicorn / PyYAML |
| 数据库 | `/opt/gpu-fan-console/app/data/fan-console.db` | **SQLite 是权威数据源**,升级时绝不能被覆盖 |
| 运行时配置 | `/opt/gpu-fan-console/app/config.yaml` | 只在**首次运行**当初始值用;之后改了不算数 |
| 运行时配置 | `/opt/gpu-fan-console/app/config.yaml` | 只在**首次运行**当初始值用;之后改了不算数。仓库里的版本是**模板**(示例地址),升级时解压要排除它(见 4.1 步骤 4) |
| 主服务 | `/etc/systemd/system/gpu-fan-console.service` | `enabled` + `active`,`Restart=always` |
| 看门狗 | `/etc/systemd/system/fan-watchdog.{service,timer}` | `enabled`,每 2 分钟查一次心跳 |
| 心跳文件 | `/run/gpu-fan-console/heartbeat` | 主进程每轮写;过期则由看门狗强推 `8×0x00` 回落 |
| 依赖的 exporter | ipmi_exporter `:9290`、node_exporter `:9100`、DCGM `:9400` | 都在 pve02 本机,都在跑 |
| 依赖的 exporter | ipmi_exporter `:9290`、node_exporter `:9100`、DCGM `:9400` | 自备组件,装法不限,见 2.1 |

### 2.1 前置依赖:三个 exporter(自备组件,装法不限)

Expand Down Expand Up @@ -72,12 +72,6 @@ curl -s localhost:9290/metrics | grep ipmi_fan_speed_rpm | head -1
curl -s localhost:9100/metrics | grep Tctl | head -1
```

> **pve02 当前实况**(仅本机维护参考,不是规定):dcgm-exporter 跑 docker
> (`nvcr.io/nvidia/k8s/dcgm-exporter:4.6.0-4.8.3-distroless`,`--restart unless-stopped`,
> `--gpus all`);ipmi_exporter v1.10.1 与 node_exporter 跑 systemd,二进制在
> `/usr/local/bin/`,node_exporter 带 `--collector.hwmon --collector.cpufreq`,
> ipmi_exporter 以 root 运行。

---

## 3. 首次部署(换机器 / 重装时才需要)
Expand All @@ -87,7 +81,7 @@ curl -s localhost:9100/metrics | grep Tctl | head -1
# 它们是控制台的全部数据源,没有这一步控速无依据

# ① 建目录、建 venv
ssh pve02
ssh <目标机>
mkdir -p /opt/gpu-fan-console
cd /opt/gpu-fan-console
python3 -m venv .venv
Expand Down Expand Up @@ -125,14 +119,17 @@ tar czf /tmp/gfc.tgz \
-C . app

# 3) 传上去
agentsshcli upload pve02 /tmp/gfc.tgz /root/gfc.tgz # 这条路现在常挂,见「坑 ②」
agentsshcli upload <连接名> /tmp/gfc.tgz /root/gfc.tgz # 这条路现在常挂,见「坑 ②」

# 4) 远端解压 + 清旧前端产物(⚠️ 见「坑 ③」)
agentsshcli exec pve02 "tar xzf /root/gfc.tgz -C /opt/gpu-fan-console"
# 4) 远端解压 + 清旧前端产物(⚠️ 见「坑 ③」)
# 已配置过的机器升级时排除 config.yaml —— 仓库里的版本是模板(示例地址),
# 别把现场配好的 prometheus_url 等覆盖回示例值。首次部署才需要带上它。
agentsshcli exec <连接名> "tar xzf /root/gfc.tgz -C /opt/gpu-fan-console --exclude='app/config.yaml'"

# 5) 重启 + 自检
agentsshcli exec pve02 "systemctl restart gpu-fan-console && sleep 4 && systemctl is-active gpu-fan-console"
curl -s http://192.168.6.7:8765/api/status | head -c 300
agentsshcli exec <连接名> "systemctl restart gpu-fan-console && sleep 4 && systemctl is-active gpu-fan-console"
curl -s http://192.0.2.10:8765/api/status | head -c 300
```

`deploy.sh` 就是把上面这套串起来,并且**内置了下面三个坑的绕过方案**。
Expand Down Expand Up @@ -188,10 +185,10 @@ split -b 30000 all.b64 ch_ # ⚠️ 单次命令行上限 32767 字符
i=0
for f in ch_*; do
[ $i -eq 0 ] && R='>' || R='>>'
agentsshcli exec pve02 "printf '%s' '$(cat $f)' $R /root/pkg.b64"
agentsshcli exec <连接名> "printf '%s' '$(cat $f)' $R /root/pkg.b64"
i=$((i+1))
done
agentsshcli exec pve02 "base64 -d /root/pkg.b64 > /root/pkg.tgz"
agentsshcli exec <连接名> "base64 -d /root/pkg.b64 > /root/pkg.tgz"
```
> 只发前端时先 `gzip -9` 再 base64,体积能砍到 1/3,分片数从 25 降到 10。

Expand All @@ -214,9 +211,9 @@ find <dir> -type f ! -name '新文件1' ! -name '新文件2' -delete
|---|---|---|
| 服务活着 | `systemctl is-active gpu-fan-console` | `active` |
| 看门狗在跑 | `systemctl list-timers fan-watchdog.timer` | 有下次触发时间 |
| 页面 200 | `curl -s -o /dev/null -w '%{http_code}' http://192.168.6.7:8765/` | `200` |
| 静态资源对得上 | `curl -s http://192.168.6.7:8765/ \| grep -o 'index-[A-Za-z0-9_-]*\.\(js\|css\)'` | 与本地 `ls app/static/assets` 一致 |
| 曲线在控速 | `curl -s http://192.168.6.7:8765/api/status` | 受控位 `duty` 有值、`temperature` 与 GPU 温度对得上 |
| 页面 200 | `curl -s -o /dev/null -w '%{http_code}' http://192.0.2.10:8765/` | `200` |
| 静态资源对得上 | `curl -s http://192.0.2.10:8765/ \| grep -o 'index-[A-Za-z0-9_-]*\.\(js\|css\)'` | 与本地 `ls app/static/assets` 一致 |
| 曲线在控速 | `curl -s http://192.0.2.10:8765/api/status` | 受控位 `duty` 有值、`temperature` 与 GPU 温度对得上 |
| 数据没丢 | 打开设置页 | 模式 / 管控 GPU / 分配关系都是部署前的样子 |
| 启动顺序对 | `journalctl -u gpu-fan-console -n 30` | 先「已应用数据库里的设置」再「已应用数据库里的分配」 |
| 回落动作在 | 同上,搜「安全护栏已激活」 | 有这行才说明退出时会回退 auto |
Expand All @@ -232,17 +229,17 @@ git checkout <上一个好提交>
./deploy.sh

# 数据回滚(SQLite 有备份时)
agentsshcli exec pve02 "systemctl stop gpu-fan-console"
agentsshcli exec pve02 "cp /root/fan-console.db.bak /opt/gpu-fan-console/app/data/fan-console.db"
agentsshcli exec pve02 "systemctl start gpu-fan-console"
agentsshcli exec <连接名> "systemctl stop gpu-fan-console"
agentsshcli exec <连接名> "cp /root/fan-console.db.bak /opt/gpu-fan-console/app/data/fan-console.db"
agentsshcli exec <连接名> "systemctl start gpu-fan-console"
```

**最坏情况的保险**:不管服务死没死,都可以直接让 BMC 接管 ——

```bash
# 远端执行(把 8 个字节全填 0x00 = 全交回 BMC 自动档)
ipmitool raw 0x3a 0x01 0x00 0x00 0x00 0x00 0x00 0x00 0x00 0x00
# 宿主都起不来时:从 BMC 独立地址 192.168.6.8 走,凭据 admin/admin
# 宿主都起不来时:从 BMC 独立管理地址走(BMC 与宿主 OS 相互独立,凭据自行保管,切勿写进公开文档)
```

---
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,7 +134,7 @@ python -m unittest discover -s app/tests -t . -v
1. **后端单测**:`python -m unittest`(Python 3.13,对齐线上)
2. **前端构建**:`npm ci && npm run build`(自带 vue-tsc 全量类型检查)
3. **部署包**:产出 `gpu-fan-console-app.tgz`(成员路径 `app/...`,排除 `app/data`),
挂在 workflow 的 Artifacts 里——下载后传到 pve02 解压重启即可,本机无需装 Node
挂在 workflow 的 Artifacts 里——下载后传到目标机解压重启即可,本机无需装 Node

推送 main 时额外构建**双平台独立可执行文件**(PyInstaller);打 `v*` tag 自动创建
GitHub Release 并附上全部产物:
Expand Down Expand Up @@ -201,7 +201,7 @@ alert: # 邮件告警(可选,默认关闭)
max_failed_attempts: 3
email: { ... } # SMTP 配置,支持多收件人、1 小时防轰炸
prometheus:
base_url: "http://192.168.6.31:30091"
base_url: "http://192.0.2.20:30091"
servers:
- type: dell730
ip: "192.168.71.90"
Expand Down
2 changes: 1 addition & 1 deletion app/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
2. **API 服务**:REST + WebSocket 给前端
3. **静态托管**:直接伺服前端构建产物

没有跨机通信、没有独立 agent —— 它只管 pve02 自己这台机器的风扇。
没有跨机通信、没有独立 agent —— 它只管部署它的这台机器的风扇。
"""

__version__ = "0.1.0"
6 changes: 3 additions & 3 deletions app/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -56,11 +56,11 @@ class SourcesConfig(BaseModel):
# 实时读数**不走**这里,因为 Prometheus 是 30s 快照(见 CLAUDE.md)。
# 但历史曲线恰恰是 Prometheus 的主场,所以单独配一套。
prometheus_url: str | None = None
#: GPU 温度在 Prometheus 里的 instance 标签,如 ``192.168.6.7:9400``
#: GPU 温度在 Prometheus 里的 instance 标签,如 ``192.0.2.10:9400``
prometheus_gpu_instance: str | None = None
#: CPU 温度(node_exporter 的 hwmon / k10temp)的 instance 标签,如 ``192.168.6.7:9100``
#: CPU 温度(node_exporter 的 hwmon / k10temp)的 instance 标签,如 ``192.0.2.10:9100``
prometheus_node_instance: str | None = None
#: 风扇转速的 instance 标签,如 ``192.168.6.7:9290``
#: 风扇转速的 instance 标签,如 ``192.0.2.10:9290``
prometheus_fan_instance: str | None = None


Expand Down
16 changes: 9 additions & 7 deletions app/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -29,21 +29,23 @@ sources:
http_timeout: 5

# 远程 BMC 兜底(宿主系统起不来时排障用)。默认走本地 in-band,不配这项。
# ⚠️ 换成你自己的 BMC 地址与凭据;凭据绝不要提交到公开仓库。
# ipmi_remote:
# host: "192.168.6.8"
# user: "admin"
# password: "admin"
# host: "192.0.2.11" # 示例地址(RFC 5737 文档段),按实际替换
# user: "your-bmc-user"
# password: "your-bmc-password"

# --- 历史趋势(可选,只给 /api/history 用)---
#
# 实时读数**不走**这里:Prometheus 是 30s 一次 scrape,拿到的永远是快照,
# 而上面两个 endpoint 是直连 exporter,读的才是当下值。
# 但历史曲线恰恰是 Prometheus 的主场,所以单独配一套。
# 不配的话界面上的趋势图会显示"暂无历史数据",其余功能不受影响。
prometheus_url: "http://192.168.6.31:30091"
prometheus_gpu_instance: "192.168.6.7:9400" # GPU 温度 ← DCGM exporter
prometheus_node_instance: "192.168.6.7:9100" # CPU 温度 ← node_exporter
prometheus_fan_instance: "192.168.6.7:9290" # 风扇转速 ← ipmi_exporter
# ⚠️ 下面的地址全部是示例(RFC 5737 文档段),按你的实际环境替换。
prometheus_url: "http://192.0.2.20:30091" # 示例:你的 Prometheus 地址
prometheus_gpu_instance: "192.0.2.10:9400" # GPU 温度 ← DCGM exporter
prometheus_node_instance: "192.0.2.10:9100" # CPU 温度 ← node_exporter
prometheus_fan_instance: "192.0.2.10:9290" # 风扇转速 ← ipmi_exporter

curve:
# 滞回带(°C):降温方向必须跌出这个带宽才降档,防止温度在阈值附近
Expand Down
8 changes: 4 additions & 4 deletions app/ipmi.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,8 @@
b1 CPU1_FAN1
b2 --(保留)
b3 REAR_FAN1
b4 REAR_FAN2 ← pve02 用于 Tesla T10 散热
b5 FRNT_FAN1 ← pve02 用于 Tesla T10 散热
b4 REAR_FAN2 ← 宿主机用于 Tesla T10 散热
b5 FRNT_FAN1 ← 宿主机用于 Tesla T10 散热
b6 FRNT_FAN2 (未接风扇)
b7 FRNT_FAN3 (未接风扇)
b8 FRNT_FAN4 (未接风扇)
Expand Down Expand Up @@ -148,9 +148,9 @@ def encode_duty(duty: int | None) -> int:
class IPMIClient:
"""ipmitool 封装。

默认走**本地 in-band**(``/dev/ipmi0``,需要 root),这是 pve02 上的推荐用法:
默认走**本地 in-band**(``/dev/ipmi0``,需要 root),这是推荐用法:
链路最短、无网络依赖。也支持 ``lanplus`` 远程模式指向 BMC 独立地址
(``192.168.6.8``),用于宿主系统起不来时的带外兜底。
(``192.0.2.11``),用于宿主系统起不来时的带外兜底。
"""

def __init__(
Expand Down
2 changes: 1 addition & 1 deletion app/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -186,7 +186,7 @@ async def lifespan(app: FastAPI):

app = FastAPI(
title="GPU Fan Console",
description="pve02 GPU 温度联动风扇控制台(单机应用)",
description="GPU 温度联动风扇控制台(单机应用)",
version=__version__,
lifespan=lifespan,
)
Expand Down
2 changes: 1 addition & 1 deletion app/sensors.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@
- 所有网络/子进程调用**带超时**。上游项目的教训:一个不带超时的阻塞读
能把整个控制线程静默挂死。

关于 pve02 这块板子(ASRock Rack EPYCD8,BMC 固件 2.20)的实测要点:
关于本机这块板子(ASRock Rack EPYCD8,BMC 固件 2.20)的实测要点:

- ``ipmitool sdr type fan`` 输出**五列**:``名称 | 传感器ID | 状态 | 阈值 | 读数``。
读数在**最后一列**。最初按「第二列是读数」写会把传感器 ID ``62h`` 当成 RPM。
Expand Down
6 changes: 3 additions & 3 deletions app/tests/test_core.py
Original file line number Diff line number Diff line change
Expand Up @@ -338,7 +338,7 @@ def test_emergency_thresholds_accept_valid_pair(self) -> None:


class TestPrometheusParsing(unittest.TestCase):
"""Prometheus 文本解析 —— 用 pve02 上抓到的真实格式。"""
"""Prometheus 文本解析 —— 用 真实设备抓到的格式。"""

SAMPLE = """\
# HELP DCGM_FI_DEV_GPU_TEMP GPU temperature (in C).
Expand Down Expand Up @@ -428,7 +428,7 @@ def test_missing_clock_stays_none(self) -> None:
class TestSdrFanParsing(unittest.TestCase):
"""``ipmitool sdr type fan`` 输出解析。

样例取自 2026-09-28 在 pve02 上的真实输出 —— 这组用例的由来就是一个
样例取自 2026-09-28 在真实设备上的输出 —— 这组用例的由来就是一个
真实 bug:最初以为「第二列是读数」,结果把传感器 ID ``62h`` 当成了 RPM。
"""

Expand Down Expand Up @@ -499,7 +499,7 @@ def test_no_reading_temperature(self) -> None:
class TestFanReaderIpmitoolFallback(unittest.TestCase):
"""ipmitool 兜底路径 —— 必须跳过未接的风扇位。

样例是 2026-09-28 在 pve02 上抓的真实输出:14 个风扇传感器位里只有
样例是 2026-09-28 在真实设备上抓的输出:14 个风扇传感器位里只有
4 个有读数,其余全是 ``No Reading``。不跳过的话界面上会凭空多出
10 个空风扇位。
"""
Expand Down
13 changes: 8 additions & 5 deletions deploy.sh
Original file line number Diff line number Diff line change
@@ -1,24 +1,24 @@
#!/usr/bin/env bash
#
# 一键部署 GPU 风扇控制台到 pve02
# 一键部署 GPU 风扇控制台到目标机
#
# ./deploy.sh 全量(后端 + 前端)→ 覆盖 → 重启 → 自检
# ./deploy.sh --static-only 只更新前端(不改后端,不重启,静态文件即时生效)
# ./deploy.sh --no-build 跳过前端构建(复用 app/static 里已有的产物)
# ./deploy.sh --no-restart 传完不重启(全量模式下慎用)
#
# 可用环境变量覆盖默认值:
# CONN=pve02 REMOTE_DIR=/opt/gpu-fan-console SERVICE=gpu-fan-console URL=http://192.168.6.7:8765
# 可用环境变量(CONN 必填,其余有默认值):
# CONN=<ssh连接名> REMOTE_DIR=/opt/gpu-fan-console SERVICE=gpu-fan-console URL=http://<目标机>:8765
#
# 为什么不用 agentsshcli upload:它现在报「创建远端续传元数据失败」,虽然数据传完了
# 但整体返回失败(--no-cache 直连模式也无效,已实测)。所以走分片 base64,见 DEPLOY.md 坑 ②。
#
set -euo pipefail

CONN="${CONN:-pve02}"
CONN="${CONN:?用法: CONN=<ssh连接名> ./deploy.sh}"
REMOTE_DIR="${REMOTE_DIR:-/opt/gpu-fan-console}"
SERVICE="${SERVICE:-gpu-fan-console}"
URL="${URL:-http://192.168.6.7:8765}"
URL="${URL:-http://127.0.0.1:8765}"
# 单次命令行上限 32767 字符(Windows),留出余量
CHUNK_SIZE="${CHUNK_SIZE:-30000}"

Expand Down Expand Up @@ -71,8 +71,11 @@ if [ "$STATIC_ONLY" -eq 1 ]; then
tar czf "$PKG" -C . app/static
else
# ⚠️ --exclude='app/data' 是在保命:那是 SQLite 权威数据源,被覆盖等于丢失全部配置
# ⚠️ --exclude='app/config.yaml':仓库里的是模板(示例地址),覆盖会毁掉现场配好的
# prometheus_url 等本机值。首次部署请手动传一次完整 config.yaml 再改。
tar czf "$PKG" \
--exclude='__pycache__' --exclude='*.pyc' --exclude='app/data' \
--exclude='app/config.yaml' \
-C . app
fi
echo " 包大小:$(wc -c <"$PKG") bytes"
Expand Down
Loading
Loading