Offline Python tools that turn a 1-bit export - an ONNX export or a Q1_0 GGUF - into a bitgpu model and produce the reference data the verification page checks against. Nothing here ships to the browser.
pip install numpy onnx onnxruntime tokenizers # GGUF conversion needs only numpyGather a work dir with the export's config.json, tokenizer.json, the .onnx graph and its
external data file (for Bonsai: onnx/model_q1.onnx + onnx/model_q1.onnx_data from
onnx-community/Bonsai-1.7B-ONNX). Then:
# 1. manifest.json + aux: the two small files the engine loads (see docs/FORMAT.md)
python tools/convert-onnx.py --model <dir> [--aux-name bonsai.aux.bin] [--ref-url <f16 safetensors url>]
# 2. golden logits from the ORIGINAL export (onnxruntime CPU), the ground truth
python tools/golden.py --model <dir>
# 3. numpy forward pass straight from the manifest, checked against the golden logits;
# --dump writes the fixtures examples/verify.html reads
python tools/reference.py --model <dir> --dump test-fixtures/forward-<tag>For plain consumption the offline step is optional: bitgpu/gguf's fromGguf(url) performs
this exact conversion in the browser from the GGUF header alone (npm run test:gguf gates
the two paths deep-equal). Run the converter when you want committed/CDN-hosted manifest
files, or fixtures for the GPU gate:
One file in, two small files out - the GGUF itself stays the data file, streamed byte-for-byte from wherever it is hosted:
# 1. manifest.json (v2, q1_0 containers) + aux (just the two LUTs, ~1.5 KB)
python tools/convert-gguf.py --gguf <dir>/Bonsai-8B-Q1_0.gguf [--ref-url <f16 safetensors url>]
# 2. there is no onnxruntime oracle for a GGUF; generate the fixtures directly from the
# manifest (--ids reuses another fixture set's prompt so checkpoints stay comparable)
python tools/reference.py --model <dir> --ids test-fixtures/forward-8b/ids.i32.bin \
--dump test-fixtures/forward-<tag>GGUF repos usually carry no tokenizer.json; take it from the model's source repo on the Hub
(the chat layer needs tokenizer.json + tokenizer_config.json next to the manifest, or
explicit URLs).
The Bonsai-27B family is the hybrid qwen3_5 arch (3:1 gated-DeltaNet linear attention +
gated full attention). ONNX Runtime cannot express the gated delta rule, so the oracle here is
HF transformers itself. tools/qwen35_numpy.py is a clean-room numpy forward of the whole
backbone (the reference the WGSL kernels must match); tools/golden-qwen35.py runs the HF model
deterministically (fp32, CPU, pure-PyTorch DeltaNet fallback - do NOT install
flash-linear-attention/causal-conv1d) and validates the numpy math layer-by-layer against it.
pip install "transformers>=5.14" torch safetensors # heavy, dev-only oracle
# validate the clean-room math against HF on the small architectural twin (both delta paths)
python tools/golden-qwen35.py --delta chunk # chunk-parallel scan == HF prefill
python tools/golden-qwen35.py --delta recurrent # sequential scan == bitgpu decodeQwen/Qwen3.5-0.8B is the small twin of the 1-bit Bonsai-27B (same arch, 24 layers) - it
runs on a laptop and is the fast dev/validation target. The 1-bit fixtures for the actual 27B
are produced from its manifest with the same numpy math (dequantized weights in, golden out).
The hybrid-specific WGSL kernels (conv1d, gated-DeltaNet scan, gated RMSNorm, g/beta, partial
RoPE, head_dim-256 attention) are validated in isolation against the numpy oracle before the
end-to-end gate: npm run test:kernels (needs numpy + system Chrome/WebGPU, a local GPU gate).
tools/gen-kernel-cases.py emits the cases; scripts/verify-kernels.mjs runs the real shaders.
After the dump, point examples/model-<tag> at the work dir (the reference Bonsai-1.7B is tag
1.7b) and run npm run verify:headless (it serves the repo itself): the engine must
reproduce the reference forward (cosine) and, for the committed Bonsai fixtures, the exact
known-good token ids (record the engine's greedy continuation as known_good in the new set's
params.json on first run).
The shipped Bonsai manifests + aux files were reproduced byte-for-byte from the public exports with these tools, so they are the authoritative producers of the format.