Skip to content

Commit f62427a

Browse files
authored
Merge pull request #8 from chennest/feature/monitoring-dashboard
安全脱敏:清除内网拓扑与凭据信息
2 parents 5e98dcf + b9c5a62 commit f62427a

16 files changed

Lines changed: 84 additions & 79 deletions

File tree

‎.github/workflows/ci.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -23,7 +23,7 @@ jobs:
2323

2424
- uses: actions/setup-python@v5
2525
with:
26-
python-version: "3.13" # 对齐 pve02 线上版本
26+
python-version: "3.13" # 对齐线上生产版本
2727

2828
- name: 后端单元测试
2929
run: |

‎CLAUDE.md‎

Lines changed: 9 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -166,22 +166,22 @@ cat logs/fancontroller.log.YYYY-MM-DD
166166

167167
```yaml
168168
prometheus:
169-
base_url: "http://192.168.6.31:30091"
169+
base_url: "http://192.0.2.20:30091"
170170
timeout: 10
171171

172172
servers:
173173
- type: epycd8
174174
prometheus: # per-server 覆盖(多机场景必需)
175-
fan_instance: "192.168.6.7:9290" # 风扇转速 ← ipmi_exporter
176-
temp_instance: "192.168.6.7:9100" # CPU 温度 ← node_exporter
177-
gpu_instance: "192.168.6.7:9400" # GPU 温度 ← DCGM exporter
175+
fan_instance: "192.0.2.10:9290" # 风扇转速 ← ipmi_exporter
176+
temp_instance: "192.0.2.10:9100" # CPU 温度 ← node_exporter
177+
gpu_instance: "192.0.2.10:9400" # GPU 温度 ← DCGM exporter
178178
```
179179
180-
| 数据 | 指标 | instance(pve02 实测) |
180+
| 数据 | 指标 | instance(实测示例) |
181181
|------|------|----------------------|
182-
| 风扇转速 | `ipmi_fan_speed_rpm{name="FRNT_FAN1"}` | `192.168.6.7:9290` |
183-
| CPU 温度 | `node_hwmon_temp_celsius`(按语义标签 `label="Tctl"` 过滤) | `192.168.6.7:9100` |
184-
| GPU 温度 | `DCGM_FI_DEV_GPU_TEMP` | `192.168.6.7:9400` |
182+
| 风扇转速 | `ipmi_fan_speed_rpm{name="FRNT_FAN1"}` | `192.0.2.10:9290` |
183+
| CPU 温度 | `node_hwmon_temp_celsius`(按语义标签 `label="Tctl"` 过滤) | `192.0.2.10:9100` |
184+
| GPU 温度 | `DCGM_FI_DEV_GPU_TEMP` | `192.0.2.10:9400` |
185185

186186
⚠️ **三个 instance 对应三个不同的 exporter,混用会直接查不到数据。** 配置项
187187
分别为 `fan_instance` / `temp_instance` / `gpu_instance`(笼统的 `instance` 仍
@@ -219,7 +219,7 @@ servers:
219219

220220
本仓库真正在用的是 `app/`(FastAPI 控制台)+ `frontend/`(Vue3 前端):
221221

222-
- 控速依据是 **GPU 温度**(DCGM),目标机 pve02 的机箱风扇,写 `ipmitool raw 0x3a 0x01`
222+
- 控速依据是 **GPU 温度**(DCGM),目标机的机箱风扇,写 `ipmitool raw 0x3a 0x01`
223223
- 运行时状态(模式 / 管控 GPU / 分配关系 / 审计)**全部落 SQLite**,配置文件只是首次运行的种子
224224

225225
**实时数据源:三个外部 exporter(自备组件,不随仓库提供,装法不限)**

‎DEPLOY.md‎

Lines changed: 22 additions & 25 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,9 @@
11
# 部署手册 —— GPU 风扇控制台
22

3-
> 目标机:**pve02**(ASRock Rack EPYCD8,192.168.6.7,BMC 2.20)
3+
> 目标机:你的 GPU 宿主机(下文 `<目标机>` / `<连接名>` 按实际替换)
44
> 部署目录:`/opt/gpu-fan-console` | 服务名:`gpu-fan-console.service`
5-
> 访问地址:`http://192.168.6.7:8765`
6-
> 最后核对:2026-09-28(当时线上 Python 3.13.5 / fastapi 0.141.1 / uvicorn 0.53.0)
5+
> 访问地址:`http://<目标机>:8765`
6+
> 参考环境:Python 3.13.5 / fastapi 0.141.1 / uvicorn 0.53.0(2026-09-28 核对)
77
88
---
99

@@ -37,11 +37,11 @@ FastAPI 挂 `StaticFiles` 一起发出去。所以:
3737
| 代码目录 | `/opt/gpu-fan-console` | 工作目录,service 的 `WorkingDirectory` |
3838
| 虚拟环境 | `/opt/gpu-fan-console/.venv` | Python 3.13.5,已装 fastapi / uvicorn / PyYAML |
3939
| 数据库 | `/opt/gpu-fan-console/app/data/fan-console.db` | **SQLite 是权威数据源**,升级时绝不能被覆盖 |
40-
| 运行时配置 | `/opt/gpu-fan-console/app/config.yaml` | 只在**首次运行**当初始值用;之后改了不算数 |
40+
| 运行时配置 | `/opt/gpu-fan-console/app/config.yaml` | 只在**首次运行**当初始值用;之后改了不算数。仓库里的版本是**模板**(示例地址),升级时解压要排除它(见 4.1 步骤 4) |
4141
| 主服务 | `/etc/systemd/system/gpu-fan-console.service` | `enabled` + `active`,`Restart=always` |
4242
| 看门狗 | `/etc/systemd/system/fan-watchdog.{service,timer}` | `enabled`,每 2 分钟查一次心跳 |
4343
| 心跳文件 | `/run/gpu-fan-console/heartbeat` | 主进程每轮写;过期则由看门狗强推 `8×0x00` 回落 |
44-
| 依赖的 exporter | ipmi_exporter `:9290`、node_exporter `:9100`、DCGM `:9400` | 都在 pve02 本机,都在跑 |
44+
| 依赖的 exporter | ipmi_exporter `:9290`、node_exporter `:9100`、DCGM `:9400` | 自备组件,装法不限,见 2.1 |
4545

4646
### 2.1 前置依赖:三个 exporter(自备组件,装法不限)
4747

@@ -72,12 +72,6 @@ curl -s localhost:9290/metrics | grep ipmi_fan_speed_rpm | head -1
7272
curl -s localhost:9100/metrics | grep Tctl | head -1
7373
```
7474

75-
> **pve02 当前实况**(仅本机维护参考,不是规定):dcgm-exporter 跑 docker
76-
> (`nvcr.io/nvidia/k8s/dcgm-exporter:4.6.0-4.8.3-distroless`,`--restart unless-stopped`,
77-
> `--gpus all`);ipmi_exporter v1.10.1 与 node_exporter 跑 systemd,二进制在
78-
> `/usr/local/bin/`,node_exporter 带 `--collector.hwmon --collector.cpufreq`,
79-
> ipmi_exporter 以 root 运行。
80-
8175
---
8276

8377
## 3. 首次部署(换机器 / 重装时才需要)
@@ -87,7 +81,7 @@ curl -s localhost:9100/metrics | grep Tctl | head -1
8781
# 它们是控制台的全部数据源,没有这一步控速无依据
8882

8983
# ① 建目录、建 venv
90-
ssh pve02
84+
ssh <目标机>
9185
mkdir -p /opt/gpu-fan-console
9286
cd /opt/gpu-fan-console
9387
python3 -m venv .venv
@@ -125,14 +119,17 @@ tar czf /tmp/gfc.tgz \
125119
-C . app
126120

127121
# 3) 传上去
128-
agentsshcli upload pve02 /tmp/gfc.tgz /root/gfc.tgz # 这条路现在常挂,见「坑 ②」
122+
agentsshcli upload <连接名> /tmp/gfc.tgz /root/gfc.tgz # 这条路现在常挂,见「坑 ②」
129123

130124
# 4) 远端解压 + 清旧前端产物(⚠️ 见「坑 ③」)
131-
agentsshcli exec pve02 "tar xzf /root/gfc.tgz -C /opt/gpu-fan-console"
125+
# 4) 远端解压 + 清旧前端产物(⚠️ 见「坑 ③」)
126+
# 已配置过的机器升级时排除 config.yaml —— 仓库里的版本是模板(示例地址),
127+
# 别把现场配好的 prometheus_url 等覆盖回示例值。首次部署才需要带上它。
128+
agentsshcli exec <连接名> "tar xzf /root/gfc.tgz -C /opt/gpu-fan-console --exclude='app/config.yaml'"
132129

133130
# 5) 重启 + 自检
134-
agentsshcli exec pve02 "systemctl restart gpu-fan-console && sleep 4 && systemctl is-active gpu-fan-console"
135-
curl -s http://192.168.6.7:8765/api/status | head -c 300
131+
agentsshcli exec <连接名> "systemctl restart gpu-fan-console && sleep 4 && systemctl is-active gpu-fan-console"
132+
curl -s http://192.0.2.10:8765/api/status | head -c 300
136133
```
137134

138135
`deploy.sh` 就是把上面这套串起来,并且**内置了下面三个坑的绕过方案**。
@@ -188,10 +185,10 @@ split -b 30000 all.b64 ch_ # ⚠️ 单次命令行上限 32767 字符
188185
i=0
189186
for f in ch_*; do
190187
[ $i -eq 0 ] && R='>' || R='>>'
191-
agentsshcli exec pve02 "printf '%s' '$(cat $f)' $R /root/pkg.b64"
188+
agentsshcli exec <连接名> "printf '%s' '$(cat $f)' $R /root/pkg.b64"
192189
i=$((i+1))
193190
done
194-
agentsshcli exec pve02 "base64 -d /root/pkg.b64 > /root/pkg.tgz"
191+
agentsshcli exec <连接名> "base64 -d /root/pkg.b64 > /root/pkg.tgz"
195192
```
196193
> 只发前端时先 `gzip -9` 再 base64,体积能砍到 1/3,分片数从 25 降到 10。
197194
@@ -214,9 +211,9 @@ find <dir> -type f ! -name '新文件1' ! -name '新文件2' -delete
214211
|---|---|---|
215212
| 服务活着 | `systemctl is-active gpu-fan-console` | `active` |
216213
| 看门狗在跑 | `systemctl list-timers fan-watchdog.timer` | 有下次触发时间 |
217-
| 页面 200 | `curl -s -o /dev/null -w '%{http_code}' http://192.168.6.7:8765/` | `200` |
218-
| 静态资源对得上 | `curl -s http://192.168.6.7:8765/ \| grep -o 'index-[A-Za-z0-9_-]*\.\(js\|css\)'` | 与本地 `ls app/static/assets` 一致 |
219-
| 曲线在控速 | `curl -s http://192.168.6.7:8765/api/status` | 受控位 `duty` 有值、`temperature` 与 GPU 温度对得上 |
214+
| 页面 200 | `curl -s -o /dev/null -w '%{http_code}' http://192.0.2.10:8765/` | `200` |
215+
| 静态资源对得上 | `curl -s http://192.0.2.10:8765/ \| grep -o 'index-[A-Za-z0-9_-]*\.\(js\|css\)'` | 与本地 `ls app/static/assets` 一致 |
216+
| 曲线在控速 | `curl -s http://192.0.2.10:8765/api/status` | 受控位 `duty` 有值、`temperature` 与 GPU 温度对得上 |
220217
| 数据没丢 | 打开设置页 | 模式 / 管控 GPU / 分配关系都是部署前的样子 |
221218
| 启动顺序对 | `journalctl -u gpu-fan-console -n 30` | 先「已应用数据库里的设置」再「已应用数据库里的分配」 |
222219
| 回落动作在 | 同上,搜「安全护栏已激活」 | 有这行才说明退出时会回退 auto |
@@ -232,17 +229,17 @@ git checkout <上一个好提交>
232229
./deploy.sh
233230

234231
# 数据回滚(SQLite 有备份时)
235-
agentsshcli exec pve02 "systemctl stop gpu-fan-console"
236-
agentsshcli exec pve02 "cp /root/fan-console.db.bak /opt/gpu-fan-console/app/data/fan-console.db"
237-
agentsshcli exec pve02 "systemctl start gpu-fan-console"
232+
agentsshcli exec <连接名> "systemctl stop gpu-fan-console"
233+
agentsshcli exec <连接名> "cp /root/fan-console.db.bak /opt/gpu-fan-console/app/data/fan-console.db"
234+
agentsshcli exec <连接名> "systemctl start gpu-fan-console"
238235
```
239236

240237
**最坏情况的保险**:不管服务死没死,都可以直接让 BMC 接管 ——
241238

242239
```bash
243240
# 远端执行(把 8 个字节全填 0x00 = 全交回 BMC 自动档)
244241
ipmitool raw 0x3a 0x01 0x00 0x00 0x00 0x00 0x00 0x00 0x00 0x00
245-
# 宿主都起不来时:从 BMC 独立地址 192.168.6.8 走,凭据 admin/admin
242+
# 宿主都起不来时:从 BMC 独立管理地址走(BMC 与宿主 OS 相互独立,凭据自行保管,切勿写进公开文档)
246243
```
247244

248245
---

‎README.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -134,7 +134,7 @@ python -m unittest discover -s app/tests -t . -v
134134
1. **后端单测**:`python -m unittest`(Python 3.13,对齐线上)
135135
2. **前端构建**:`npm ci && npm run build`(自带 vue-tsc 全量类型检查)
136136
3. **部署包**:产出 `gpu-fan-console-app.tgz`(成员路径 `app/...`,排除 `app/data`),
137-
挂在 workflow 的 Artifacts 里——下载后传到 pve02 解压重启即可,本机无需装 Node
137+
挂在 workflow 的 Artifacts 里——下载后传到目标机解压重启即可,本机无需装 Node
138138

139139
推送 main 时额外构建**双平台独立可执行文件**(PyInstaller);打 `v*` tag 自动创建
140140
GitHub Release 并附上全部产物:
@@ -201,7 +201,7 @@ alert: # 邮件告警(可选,默认关闭)
201201
max_failed_attempts: 3
202202
email: { ... } # SMTP 配置,支持多收件人、1 小时防轰炸
203203
prometheus:
204-
base_url: "http://192.168.6.31:30091"
204+
base_url: "http://192.0.2.20:30091"
205205
servers:
206206
- type: dell730
207207
ip: "192.168.71.90"

‎app/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -5,7 +5,7 @@
55
2. **API 服务**:REST + WebSocket 给前端
66
3. **静态托管**:直接伺服前端构建产物
77
8-
没有跨机通信、没有独立 agent —— 它只管 pve02 自己这台机器的风扇。
8+
没有跨机通信、没有独立 agent —— 它只管部署它的这台机器的风扇。
99
"""
1010

1111
__version__ = "0.1.0"

‎app/config.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -56,11 +56,11 @@ class SourcesConfig(BaseModel):
5656
# 实时读数**不走**这里,因为 Prometheus 是 30s 快照(见 CLAUDE.md)。
5757
# 但历史曲线恰恰是 Prometheus 的主场,所以单独配一套。
5858
prometheus_url: str | None = None
59-
#: GPU 温度在 Prometheus 里的 instance 标签,如 ``192.168.6.7:9400``
59+
#: GPU 温度在 Prometheus 里的 instance 标签,如 ``192.0.2.10:9400``
6060
prometheus_gpu_instance: str | None = None
61-
#: CPU 温度(node_exporter 的 hwmon / k10temp)的 instance 标签,如 ``192.168.6.7:9100``
61+
#: CPU 温度(node_exporter 的 hwmon / k10temp)的 instance 标签,如 ``192.0.2.10:9100``
6262
prometheus_node_instance: str | None = None
63-
#: 风扇转速的 instance 标签,如 ``192.168.6.7:9290``
63+
#: 风扇转速的 instance 标签,如 ``192.0.2.10:9290``
6464
prometheus_fan_instance: str | None = None
6565

6666

‎app/config.yaml‎

Lines changed: 9 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -29,21 +29,23 @@ sources:
2929
http_timeout: 5
3030

3131
# 远程 BMC 兜底(宿主系统起不来时排障用)。默认走本地 in-band,不配这项。
32+
# ⚠️ 换成你自己的 BMC 地址与凭据;凭据绝不要提交到公开仓库。
3233
# ipmi_remote:
33-
# host: "192.168.6.8"
34-
# user: "admin"
35-
# password: "admin"
34+
# host: "192.0.2.11" # 示例地址(RFC 5737 文档段),按实际替换
35+
# user: "your-bmc-user"
36+
# password: "your-bmc-password"
3637

3738
# --- 历史趋势(可选,只给 /api/history 用)---
3839
#
3940
# 实时读数**不走**这里:Prometheus 是 30s 一次 scrape,拿到的永远是快照,
4041
# 而上面两个 endpoint 是直连 exporter,读的才是当下值。
4142
# 但历史曲线恰恰是 Prometheus 的主场,所以单独配一套。
4243
# 不配的话界面上的趋势图会显示"暂无历史数据",其余功能不受影响。
43-
prometheus_url: "http://192.168.6.31:30091"
44-
prometheus_gpu_instance: "192.168.6.7:9400" # GPU 温度 ← DCGM exporter
45-
prometheus_node_instance: "192.168.6.7:9100" # CPU 温度 ← node_exporter
46-
prometheus_fan_instance: "192.168.6.7:9290" # 风扇转速 ← ipmi_exporter
44+
# ⚠️ 下面的地址全部是示例(RFC 5737 文档段),按你的实际环境替换。
45+
prometheus_url: "http://192.0.2.20:30091" # 示例:你的 Prometheus 地址
46+
prometheus_gpu_instance: "192.0.2.10:9400" # GPU 温度 ← DCGM exporter
47+
prometheus_node_instance: "192.0.2.10:9100" # CPU 温度 ← node_exporter
48+
prometheus_fan_instance: "192.0.2.10:9290" # 风扇转速 ← ipmi_exporter
4749

4850
curve:
4951
# 滞回带(°C):降温方向必须跌出这个带宽才降档,防止温度在阈值附近

‎app/ipmi.py‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -23,8 +23,8 @@
2323
b1 CPU1_FAN1
2424
b2 --(保留)
2525
b3 REAR_FAN1
26-
b4 REAR_FAN2 ← pve02 用于 Tesla T10 散热
27-
b5 FRNT_FAN1 ← pve02 用于 Tesla T10 散热
26+
b4 REAR_FAN2 ← 宿主机用于 Tesla T10 散热
27+
b5 FRNT_FAN1 ← 宿主机用于 Tesla T10 散热
2828
b6 FRNT_FAN2 (未接风扇)
2929
b7 FRNT_FAN3 (未接风扇)
3030
b8 FRNT_FAN4 (未接风扇)
@@ -148,9 +148,9 @@ def encode_duty(duty: int | None) -> int:
148148
class IPMIClient:
149149
"""ipmitool 封装。
150150
151-
默认走**本地 in-band**(``/dev/ipmi0``,需要 root),这是 pve02 上的推荐用法:
151+
默认走**本地 in-band**(``/dev/ipmi0``,需要 root),这是推荐用法:
152152
链路最短、无网络依赖。也支持 ``lanplus`` 远程模式指向 BMC 独立地址
153-
(``192.168.6.8``),用于宿主系统起不来时的带外兜底。
153+
(``192.0.2.11``),用于宿主系统起不来时的带外兜底。
154154
"""
155155

156156
def __init__(

‎app/main.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -186,7 +186,7 @@ async def lifespan(app: FastAPI):
186186

187187
app = FastAPI(
188188
title="GPU Fan Console",
189-
description="pve02 GPU 温度联动风扇控制台(单机应用)",
189+
description="GPU 温度联动风扇控制台(单机应用)",
190190
version=__version__,
191191
lifespan=lifespan,
192192
)

‎app/sensors.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,7 @@
1313
- 所有网络/子进程调用**带超时**。上游项目的教训:一个不带超时的阻塞读
1414
能把整个控制线程静默挂死。
1515
16-
关于 pve02 这块板子(ASRock Rack EPYCD8,BMC 固件 2.20)的实测要点:
16+
关于本机这块板子(ASRock Rack EPYCD8,BMC 固件 2.20)的实测要点:
1717
1818
- ``ipmitool sdr type fan`` 输出**五列**:``名称 | 传感器ID | 状态 | 阈值 | 读数``。
1919
读数在**最后一列**。最初按「第二列是读数」写会把传感器 ID ``62h`` 当成 RPM。

0 commit comments

Comments
 (0)