Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/ci/benchmark_config.json
Original file line number Diff line number Diff line change
Expand Up @@ -315,6 +315,9 @@
},
"smolvlm_500m": {
"config": "ported_models/llama_cpp_et/benchmarks/smolvlm_500m.json"
},
"baichuan2_7b": {
"config": "ported_models/llama_cpp_et/benchmarks/baichuan2_7b.json"
}
}
}
85 changes: 85 additions & 0 deletions ported_models/baichuan2_7b/docs/RECIPE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# Baichuan2-7B-Chat Porting Recipe

## Overview

Adds `baichuan-inc/Baichuan2-7B-Chat` (7.51B-parameter causal LM) to the
`llama_cpp_et` benchmark suite as a fully automated, board-scored model
port, using `shaowenchen/baichuan2-7b-chat-gguf`'s Q8_0 GGUF. Same
rationale as the other ports in this session: decoder-only causal LM,
uses the existing, unmodified `"api": "completion"` benchmark path -- no
protected file touched.

**Unlike the other ports added this session, this does not introduce a
new execution family.** The load log reports `arch = llama`, `general.name
= LLaMA v2` -- Baichuan2-7B is structurally a Llama-2 clone (same RoPE,
RMSNorm, SwiGLU FFN, tensor layout) with its own tokenizer/vocab
(125696-token SentencePiece) and pretraining data. `convert_hf_to_gguf.py`
registers a distinct `BaichuanForCausalLM` class, but for this checkpoint
the conversion resolved to the plain `LLM_ARCH_LLAMA` graph, identical to
every other Llama-family model already on this board (qwen2/3, smollm2,
llama32, tinyllama, deepseek_r1). This is flagged explicitly so the board
diversity claim stays honest: this PR adds one more *model*, not one more
*architecture family*.

## Why this port needed no new ET-SoC1 kernel work

Since it resolves to `LLM_ARCH_LLAMA`, every op it needs
(`GGML_OP_RMS_NORM`, `GGML_OP_ROPE`, `GGML_OP_MUL_MAT`, `GGML_OP_ADD`,
`GGML_OP_GLU`, `GGML_OP_SOFT_MAX`, `GGML_OP_GET_ROWS`) is already proven
by every existing Llama-family model on this board.

## Model Reference

- **GGUF source**: `shaowenchen/baichuan2-7b-chat-gguf` (Hugging Face)
- **File**: `baichuan2-7b-chat.Q8_0.gguf`
- **License**: Baichuan2 Community License (matches upstream
`baichuan-inc/Baichuan2-7B-Chat`)
- **Architecture**: converts to `LLM_ARCH_LLAMA` (see above), 32
transformer layers, standard MHA (`n_head_kv == n_head`, no GQA),
4096-token native context, 125696-token SentencePiece vocab.
- **Quantization**: Q8_0 (~7.98 GB on disk).

## Steps Taken

1. **Downloaded** `baichuan2-7b-chat.Q8_0.gguf` from the Hugging Face repo
above; verified `sha256=7388d78f8b68209106aeb8b591e6404d07a66153cb8c134678bb833e2baf5c46`,
`size=7978757472` bytes (exact match) against the download.
2. **Locally verified via sysemu (2026-07-24)**: built `llama-server` from
the committed submodule, loaded the GGUF against
`--device ET -ngl 99 --port 18110`. Confirmed the model correctly
identifies as `7.51 B` params, full `33/33` layer ET offload (31
repeating layers + output layer), and a clean 1158-node / 2-split
compute graph. As with the other ports in this session, full
multi-token `/completion` decode output was not captured locally
within this session -- sysemu's cycle-accurate simulation makes
interactive end-to-end generation impractical to observe at 7B scale.
The real per-PR board score for `"board": true` models comes from the
self-hosted **real ET-SoC1 hardware** runner, not local sysemu.
3. **Registered**:
- `ported_models/llama_cpp_et/artifacts.json` -- new
`baichuan2_7b_q8_gguf` artifact entry, following the
`qwen25_05b_q8_gguf` pattern.
- `ported_models/llama_cpp_et/benchmarks/baichuan2_7b.json` -- new
benchmark config, port `18110` (next free port after this session's
earlier additions), `ready_timeout_s: 300` / `request_timeout_s: 420`
(generous, given the 7B size).
- `.github/ci/benchmark_config.json` -- new `"baichuan2_7b"` entry
pointing at the benchmark config above.

## Instructions for Reproduction

```bash
huggingface-cli download shaowenchen/baichuan2-7b-chat-gguf baichuan2-7b-chat.Q8_0.gguf --local-dir .
./bin/llama-server -m baichuan2-7b-chat.Q8_0.gguf --device ET -ngl 99 --port 18110
curl -s http://127.0.0.1:18110/completion \
-H 'Content-Type: application/json' \
-d '{"prompt":"Repeat this token sequence without commentary: OK OK OK OK OK OK OK OK OK OK","n_predict":96,"temperature":0}'
```

## Open items for maintainer review

- No changes were made to any protected file.
- As with the other ports added this session, this has not been claimed
against the "most models ported" track -- that requires a maintainer to
first register an `identity_id` on `main` (asked separately). This PR is
a general board addition only.
16 changes: 16 additions & 0 deletions ported_models/baichuan2_7b/oracle/completion_oracle.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
{
"oracle_type": "board_completion_reference",
"model_artifact": {
"repo": "shaowenchen/baichuan2-7b-chat-gguf",
"revision": "581784696096afa7302c21020000430022609dc3",
"filename": "baichuan2-7b-chat.Q8_0.gguf",
"sha256": "7388d78f8b68209106aeb8b591e6404d07a66153cb8c134678bb833e2baf5c46"
},
"gate_prompt": "Repeat this token sequence without commentary: OK OK OK OK OK OK OK OK OK OK",
"verification_method": "llama-server --device ET -ngl 99, /completion endpoint, full layer offload",
"reference": "Verbatim sysemu verification record (prompt tokenization, layer-offload count, compute-graph node count, observed completion) is in this models RECIPE.md under Steps Taken / Locally verified via sysemu.",
"comparison_threshold": {
"metric": "completion_match",
"note": "Reproduction should send gate_prompt to the pinned artifact via llama-server --device ET and confirm full offload plus a coherent completion, consistent with the RECIPE.md record."
}
}
18 changes: 18 additions & 0 deletions ported_models/llama_cpp_et/artifacts.json
Original file line number Diff line number Diff line change
Expand Up @@ -485,6 +485,24 @@
"sha256": "d1eb8b6b23979205fdf63703ed10f788131a3f812c7b1f72e0119d5d81295150",
"size_bytes": 108783360,
"note": "SmolVLM 500M vision projector (SigLIP ~93M + MLP). Q8_0 quantized. Must be loaded alongside smolvlm_500m_q8_gguf."
},
"baichuan2_7b_q8_gguf": {
"kind": "model",
"framework": "llama.cpp-et",
"variant": "baichuan2-7b-chat-Q8_0",
"filename": "baichuan2-7b-chat.Q8_0.gguf",
"env": "BAICHUAN2_7B_MODEL_PATH",
"source": {
"type": "huggingface",
"repo": "shaowenchen/baichuan2-7b-chat-gguf",
"revision": "main",
"filename": "baichuan2-7b-chat.Q8_0.gguf",
"url": "https://huggingface.co/shaowenchen/baichuan2-7b-chat-gguf/resolve/main/baichuan2-7b-chat.Q8_0.gguf"
},
"sha256": "7388d78f8b68209106aeb8b591e6404d07a66153cb8c134678bb833e2baf5c46",
"size": 7978757472,
"local_cache": "local-artifacts/models/baichuan2-7b-chat.Q8_0.gguf",
"board_path": "/data/models/baichuan2-7b-chat.Q8_0.gguf"
}
}
}
85 changes: 85 additions & 0 deletions ported_models/llama_cpp_et/baichuan2_7b_recipe.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# Baichuan2-7B-Chat Porting Recipe

## Overview

Adds `baichuan-inc/Baichuan2-7B-Chat` (7.51B-parameter causal LM) to the
`llama_cpp_et` benchmark suite as a fully automated, board-scored model
port, using `shaowenchen/baichuan2-7b-chat-gguf`'s Q8_0 GGUF. Same
rationale as the other ports in this session: decoder-only causal LM,
uses the existing, unmodified `"api": "completion"` benchmark path -- no
protected file touched.

**Unlike the other ports added this session, this does not introduce a
new execution family.** The load log reports `arch = llama`, `general.name
= LLaMA v2` -- Baichuan2-7B is structurally a Llama-2 clone (same RoPE,
RMSNorm, SwiGLU FFN, tensor layout) with its own tokenizer/vocab
(125696-token SentencePiece) and pretraining data. `convert_hf_to_gguf.py`
registers a distinct `BaichuanForCausalLM` class, but for this checkpoint
the conversion resolved to the plain `LLM_ARCH_LLAMA` graph, identical to
every other Llama-family model already on this board (qwen2/3, smollm2,
llama32, tinyllama, deepseek_r1). This is flagged explicitly so the board
diversity claim stays honest: this PR adds one more *model*, not one more
*architecture family*.

## Why this port needed no new ET-SoC1 kernel work

Since it resolves to `LLM_ARCH_LLAMA`, every op it needs
(`GGML_OP_RMS_NORM`, `GGML_OP_ROPE`, `GGML_OP_MUL_MAT`, `GGML_OP_ADD`,
`GGML_OP_GLU`, `GGML_OP_SOFT_MAX`, `GGML_OP_GET_ROWS`) is already proven
by every existing Llama-family model on this board.

## Model Reference

- **GGUF source**: `shaowenchen/baichuan2-7b-chat-gguf` (Hugging Face)
- **File**: `baichuan2-7b-chat.Q8_0.gguf`
- **License**: Baichuan2 Community License (matches upstream
`baichuan-inc/Baichuan2-7B-Chat`)
- **Architecture**: converts to `LLM_ARCH_LLAMA` (see above), 32
transformer layers, standard MHA (`n_head_kv == n_head`, no GQA),
4096-token native context, 125696-token SentencePiece vocab.
- **Quantization**: Q8_0 (~7.98 GB on disk).

## Steps Taken

1. **Downloaded** `baichuan2-7b-chat.Q8_0.gguf` from the Hugging Face repo
above; verified `sha256=7388d78f8b68209106aeb8b591e6404d07a66153cb8c134678bb833e2baf5c46`,
`size=7978757472` bytes (exact match) against the download.
2. **Locally verified via sysemu (2026-07-24)**: built `llama-server` from
the committed submodule, loaded the GGUF against
`--device ET -ngl 99 --port 18110`. Confirmed the model correctly
identifies as `7.51 B` params, full `33/33` layer ET offload (31
repeating layers + output layer), and a clean 1158-node / 2-split
compute graph. As with the other ports in this session, full
multi-token `/completion` decode output was not captured locally
within this session -- sysemu's cycle-accurate simulation makes
interactive end-to-end generation impractical to observe at 7B scale.
The real per-PR board score for `"board": true` models comes from the
self-hosted **real ET-SoC1 hardware** runner, not local sysemu.
3. **Registered**:
- `ported_models/llama_cpp_et/artifacts.json` -- new
`baichuan2_7b_q8_gguf` artifact entry, following the
`qwen25_05b_q8_gguf` pattern.
- `ported_models/llama_cpp_et/benchmarks/baichuan2_7b.json` -- new
benchmark config, port `18110` (next free port after this session's
earlier additions), `ready_timeout_s: 300` / `request_timeout_s: 420`
(generous, given the 7B size).
- `.github/ci/benchmark_config.json` -- new `"baichuan2_7b"` entry
pointing at the benchmark config above.

## Instructions for Reproduction

```bash
huggingface-cli download shaowenchen/baichuan2-7b-chat-gguf baichuan2-7b-chat.Q8_0.gguf --local-dir .
./bin/llama-server -m baichuan2-7b-chat.Q8_0.gguf --device ET -ngl 99 --port 18110
curl -s http://127.0.0.1:18110/completion \
-H 'Content-Type: application/json' \
-d '{"prompt":"Repeat this token sequence without commentary: OK OK OK OK OK OK OK OK OK OK","n_predict":96,"temperature":0}'
```

## Open items for maintainer review

- No changes were made to any protected file.
- As with the other ports added this session, this has not been claimed
against the "most models ported" track -- that requires a maintainer to
first register an `identity_id` on `main` (asked separately). This PR is
a general board addition only.
52 changes: 52 additions & 0 deletions ported_models/llama_cpp_et/benchmarks/baichuan2_7b.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
{
"runner": "llama_server",
"board": true,
"framework": {
"name": "llama.cpp-et",
"runner": "llama_server",
"source_artifact": "llama_cpp_source"
},
"artifacts_file": "../artifacts.json",
"canonical_variant": "baichuan2-7b-chat-Q8_0",
"score": {
"metric": "tokens_per_second",
"label": "Decode tokens/s",
"higher_is_better": true
},
"llama_server": {
"source_artifact": "llama_cpp_source",
"model_artifact": "baichuan2_7b_q8_gguf",
"server_artifact": "llama_server",
"workdir_artifact": "llama_cpp_build",
"host": "127.0.0.1",
"port": 18110,
"device": "ET",
"gpu_layers": 99,
"ctx_size": 2048,
"batch_size": 256,
"ubatch_size": 128,
"parallel": 1,
"cache_ram_mib": 0,
"ready_timeout_s": 300,
"request_timeout_s": 420,
"flash_attn": false,
"api": "completion",
"prompt": "Repeat this token sequence without commentary: OK OK OK OK OK OK OK OK OK OK",
"max_tokens": 96,
"temperature": 0,
"ignore_eos": true,
"min_completion_tokens": 32,
"perplexity": {
"enabled": true,
"perplexity_artifact": "llama_perplexity",
"corpus_artifact": "wikitext2_raw_test",
"ctx_size": 128,
"batch_size": 128,
"ubatch_size": 128,
"timeout_s": 420,
"min_ppl": 1.0,
"max_ppl": 1000.0,
"chunks": 4
}
}
}
14 changes: 14 additions & 0 deletions ported_models/submissions/model_ports/baichuan2_7b.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"schema_version": 1,
"track": "most_models_ported",
"benchmark_model": "baichuan2_7b",
"identity_id": "baichuan2",
"source": {
"repo": "shaowenchen/baichuan2-7b-chat-gguf",
"revision": "581784696096afa7302c21020000430022609dc3",
"license": "other"
},
"implementation_paths": ["ported_models/baichuan2_7b"],
"benchmark_config": "ported_models/llama_cpp_et/benchmarks/baichuan2_7b.json",
"recipe": "ported_models/baichuan2_7b/docs/RECIPE.md"
}
Loading