Commit 802c574
authored
[Benchmark] Upgrade benchmark args for new vllm version (#3218)
### What this PR does / why we need it?
Since the newest vllm commit has deprecated the arg `--endpoint-type`,
we should use `--backend` instead
### Does this PR introduce _any_ user-facing change?
### How was this patch tested?
test it locally:
```shell
export VLLM_USE_MODELSCOPE=true
export DATASET_PATH=/root/.cache/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
vllm serve Qwen/Qwen2.5-7B-Instruct --load-format dummy
wget -O ${DATASET_PATH} /root/.cache/datasets/ShareGPT_V3_unfiltered_cleaned_split.json https://hf-mirror.com/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json
vllm bench serve --model Qwen/Qwen2.5-7B-Instruct --backend vllm --dataset-name sharegpt --dataset-path ${DATASET_PATH} --num-prompt 200
```
and the result looks good:
```shell
============ Serving Benchmark Result ============
Successful requests: 200
Benchmark duration (s): 20.36
Total input tokens: 43560
Total generated tokens: 44697
Request throughput (req/s): 9.82
Output token throughput (tok/s): 2194.88
Peak output token throughput (tok/s): 4676.00
Peak concurrent requests: 200.00
Total Token throughput (tok/s): 4333.93
---------------Time to First Token----------------
Mean TTFT (ms): 2143.85
Median TTFT (ms): 2486.17
P99 TTFT (ms): 2530.36
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 43.50
Median TPOT (ms): 30.75
P99 TPOT (ms): 309.22
---------------Inter-token Latency----------------
Mean ITL (ms): 28.15
Median ITL (ms): 25.42
P99 ITL (ms): 38.30
==================================================
```
- vLLM version: v0.11.0rc3
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.11.0
Signed-off-by: wangli <[email protected]>1 parent 1b270a6 commit 802c574
1 file changed
+3
-3
lines changed| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
18 | 18 | | |
19 | 19 | | |
20 | 20 | | |
21 | | - | |
| 21 | + | |
22 | 22 | | |
23 | 23 | | |
24 | 24 | | |
| |||
45 | 45 | | |
46 | 46 | | |
47 | 47 | | |
48 | | - | |
| 48 | + | |
49 | 49 | | |
50 | 50 | | |
51 | 51 | | |
| |||
69 | 69 | | |
70 | 70 | | |
71 | 71 | | |
72 | | - | |
| 72 | + | |
73 | 73 | | |
74 | 74 | | |
75 | 75 | | |
| |||
0 commit comments