tools/ contains artifact conversion and inspection, benchmark orchestration, and serving smoke
checks. To download and run an existing artifact, start with the project README.
To build your own weights, use the weight conversion guide.
Run commands from the repository root with a Python environment containing the dependencies for the selected tool. The maintained environment uses Python 3.11.
Python tools are independent of CMake; there is no NINFER_BUILD_TOOLS option.
| Task | Location |
|---|---|
| Convert weights with an official or custom recipe | convert/; user guide |
| Inspect artifact metadata and objects | artifact/inspect.py |
| One-time upgrade of official v2 artifacts | upgrade_ninfer_v2_to_v3.py, with positional INPUT OUTPUT paths |
| Run benchmark matrices | bench/ |
| Measure external Serve TTFT | bench/ttft/ |
| Exercise a resident HTTP server | smoke/serve_contract.py |
| Exercise thinking preservation through a managed server | smoke/serve_thinking_preservation.py |
| Measure the physical HBM read/copy ceiling | hbm_bandwidth_probe.cu; build command |
This maintainer probe has an explicit standalone CUDA build, independent of the CMake benchmark targets. Build it with the project's CUDA toolkit and run it from the repository root:
mkdir -p build
nvcc -O3 -std=c++17 -arch=sm_120a tools/hbm_bandwidth_probe.cu \
-o build/hbm_bandwidth_probe
./build/hbm_bandwidth_probeThe common converter reads selected local sources and writes a .ninfer artifact plus its
.conversion.json report. These examples include the optional weights used by the official
artifacts. The input paths are placeholders for local checkpoint checkouts:
python3 -m tools.convert \
--model /path/to/Qwen3.6-27B \
--recipe qwen3_6_27b --components text,vision,mtp --proposal \
--resource chat_template.jinja=tools/chat_templates/qwen3_6.jinja \
--name qwen3.6-27b \
--out out/qwen3_6_27b.ninfer
python3 -m tools.convert \
--model /path/to/Qwen3.8-27B \
--recipe qwen3_8_27b --components text,vision,mtp,dflash2 --proposal \
--source dflash2=/path/to/Qwen3.8-27B-DFlash2 \
--resource chat_template.jinja=tools/chat_templates/qwen3_8.jinja \
--name qwen3.8-27b \
--out out/qwen3_8_27b.ninfer
python3 -m tools.convert \
--model /path/to/Qwen3.6-35B-A3B-base \
--recipe qwen3_6_35b_a3b --components text,vision,mtp,dflash --proposal \
--source dflash=/path/to/Qwen3.6-35B-A3B-DFlash \
--resource chat_template.jinja=tools/chat_templates/qwen3_6.jinja \
--name qwen3.6-35b-a3b \
--out out/qwen3_6_35b_a3b.ninferInspect a result:
python3 -m tools.artifact.inspect out/qwen3_6_27b.ninfer --objectsRecipes, mixed sources, custom methods, resources and sharding are described in the conversion guide. Numeric formats, layouts and framing are defined by the references linked from the documentation map.
tools/bench/run_ninfer_bench_matrix.py builds and runs the public-Engine benchmark matrix and
writes ignored local reports below profiles/bench/:
python3 tools/bench/run_ninfer_bench_matrix.py --preset core --dry-run
python3 tools/bench/run_ninfer_bench_matrix.py --preset coreSee tools/bench/README.md and bench/README.md for the
orchestrator and executable contracts.
For request-arrival latency, use the managed Qwen3.8-27B NVFP4/FP8 TTFT campaign. Its measurement
runner remains an external-only HTTP client; the separate controller owns Serve lifecycle and
artifacts. See tools/bench/ttft/README.md.
After starting ninfer-serve in another terminal:
python3 -m tools.smoke.serve_contract \
--base-url http://127.0.0.1:18080 \
--model qwen3.6-27bThe client exercises OpenAI, Anthropic, streaming, usage, multimodal, and tool-call response surfaces against the resident process.
For typed rewrite-checkpoint and thinking-history behavior, the managed smoke script launches a real server and consumes the repository fixture:
python3 tools/smoke/serve_thinking_preservation.py \
--artifact out/qwen3_6_27b.ninfer --backend mtp