Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 38 additions & 0 deletions .github/workflows/linux-reproduction-docs.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
name: Linux reproduction docs

on:
pull_request:
paths:
- 'docs/**'
- 'probes/**/*.md'
- 'probes/**/*.txt'
- 'benchmarks/**/*.md'
- 'benchmarks/**/*.txt'
- 'README.md'
- 'CONTRIBUTING.md'
- 'scripts/check_linux_repro_docs.py'
- 'tests/test_linux_repro_docs.py'
- '.github/workflows/linux-reproduction-docs.yml'
workflow_dispatch:

permissions:
contents: read

jobs:
linux-docs:
runs-on: ubuntu-24.04
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065
with:
python-version: '3.12'
- name: Check public instructions and Bash syntax
run: python scripts/check_linux_repro_docs.py --check-bash
- name: Exercise invalid-input rejection
run: |
set -euo pipefail
python -m pip install pytest==9.1.1
python -m pytest tests/test_linux_repro_docs.py -q
173 changes: 166 additions & 7 deletions .github/workflows/llm-deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,12 +6,14 @@ name: LLM Deploy v1.0
# one tests the deployment project. A failure here should not read as "the
# compiler is broken".
#
# The PR job executes small numeric gates. A separate manual/nightly job
# downloads and validates the pinned full-model release with ORT.
# PRs execute small numeric gates and a separate complete frontend audit.
# Both are mandatory for the W2 acceptance result. Full-model ORT remains
# manual/nightly; frontend parsing never claims full IR execution.
#
# docs:interfaces docs:risks
# probe:qemu-matmul probe:small-transformer probe:qwen3-onnx
# numeric:ir-qwen3-small
# numeric:ir-qwen3-small frontend:parse-qwen3
# unit:frontend-ops unit:backend-ops unit:runtime integration:runtime-model
#
# Later weeks' gates (numeric:ir-full-qwen3 at W3, numeric:qemu-full-qwen3 at
# W4, e2e:qemu-paris at W5, release:gate at W8) are NOT stubbed here. They get
Expand All @@ -23,18 +25,28 @@ on:
- 'docs/llm-deploy-v1.0/**'
- 'probes/**'
- 'scripts/check_llm_deploy_docs.py'
- 'scripts/check_linux_repro_docs.py'
- 'tests/test_linux_repro_docs.py'
- 'scripts/check_w1_repro_env.py'
- 'scripts/run_w2_acceptance.py'
- 'scripts/check_w2_ci_results.py'
- 'export_qwen3_onnx.py'
- 'scratchv/**'
- 'requirements/**'
- 'pyproject.toml'
- 'tests/test_qwen3_small*.py'
- 'tests/test_qwen3_full*.py'
- 'tests/test_tensor*.py'
- 'tests/test_riscv_tensor*.py'
- 'tests/test_optimizer*.py'
- 'tests/qwen_artifacts.py'
- 'tests/test_qwen_artifact_selection.py'
- 'tests/test_w1_qwen3*.py'
- 'tests/test_w1_repro_env.py'
- 'tests/test_qwen3_frontend_patterns.py'
- 'tests/test_qwen3_tokenizer.py'
- 'tests/test_llm_inputs.py'
- 'tests/test_w2_*.py'
- '.github/workflows/llm-deploy.yml'
schedule:
# Nightly, off the hour: 19:23 UTC = 03:23 北京时间.
Expand All @@ -50,12 +62,13 @@ env:
PROBE_TRANSFORMER: probes/w1_tiny_transformer/run.py
PROBE_QWEN3_EXPORT: probes/w1_qwen3_export/run.py
PROBE_QWEN3_SMALL: probes/w2_qwen3_small/run.py
PROBE_QWEN3_PARSE: probes/w2_qwen3_parse/run.py
QWEN3_PY: output/qwen3-probe-venv/bin/python

jobs:
llm-deploy:
name: "LLM Deploy v1.0 / W1"
runs-on: ubuntu-latest
name: "LLM Deploy v1.0 / W1 + W2"
runs-on: ubuntu-24.04
timeout-minutes: 60
steps:
- uses: actions/checkout@v4
Expand Down Expand Up @@ -99,6 +112,9 @@ jobs:
id: docs
run: python3 scripts/check_llm_deploy_docs.py --week W1

- name: "docs:linux-reproduction"
run: python3 scripts/check_linux_repro_docs.py --check-bash

# ── probe:qemu-matmul ─────────────────────────────────────────────
- name: "probe:qemu-matmul"
id: matmul
Expand All @@ -108,7 +124,15 @@ jobs:

- name: Tensor compiler and runtime regressions
run: |
python3 -m pytest tests/test_tensor_c_codegen.py tests/test_tensor_compiler.py tests/test_riscv_tensor_runtime.py tests/test_optimizer_numeric_semantics.py tests/test_qwen_artifact_selection.py tests/test_w1_qwen3_export.py tests/test_qwen3_small_riscv_gate.py tests/test_w1_repro_env.py -q
python3 -m pytest \
tests/test_tensor_c_codegen.py tests/test_tensor_compiler.py \
tests/test_riscv_tensor_runtime.py tests/test_optimizer_numeric_semantics.py \
tests/test_qwen_artifact_selection.py tests/test_w1_qwen3_export.py tests/test_qwen3_small_riscv_gate.py \
tests/test_w1_repro_env.py tests/test_qwen3_full_structure.py \
tests/test_qwen3_full_audit.py tests/test_qwen3_full_parse.py \
tests/test_llm_inputs.py tests/test_w2_backend_ops_gate.py \
tests/test_w2_runtime_gate.py tests/test_w2_runtime_metadata.py \
tests/test_w2_runtime_reporting.py tests/test_w2_acceptance.py tests/test_w2_ci_results.py -q

- name: Repeated Linux timeout cleanup regression
timeout-minutes: 5
Expand All @@ -123,6 +147,18 @@ jobs:
--junit-xml="output/process-cleanup-repeats/attempt-$attempt.xml"
done

- name: "unit:frontend-ops"
id: frontend_ops
run: |
python3 -m pytest tests/test_qwen3_frontend_patterns.py -q \
--junit-xml=output/w2-frontend-ops.xml

- name: "unit:backend-ops"
id: backend_ops
timeout-minutes: 15
run: |
python3 probes/w2_backend_ops/run.py --output-dir output/w2-backend-ops

# ── probe:small-transformer ───────────────────────────────────────
- name: "probe:small-transformer"
id: transformer
Expand Down Expand Up @@ -151,6 +187,34 @@ jobs:
"$QWEN3_PY" -X utf8 scripts/check_w1_repro_env.py \
--require-clean --output-dir output/w1-repro-env

- name: Runtime unit regressions in pinned environment
run: |
"$QWEN3_PY" -m pytest \
tests/test_llm_inputs.py tests/test_qwen3_tokenizer.py \
tests/test_w2_runtime_gate.py tests/test_w2_runtime_metadata.py \
tests/test_w2_runtime_reporting.py tests/test_w2_runtime_model.py \
-q --junit-xml=output/w2-runtime-tests.xml

- name: "unit:runtime"
id: runtime
timeout-minutes: 10
run: |
"$QWEN3_PY" probes/w2_runtime/run.py --mode download \
--tokenizer-dir output/qwen3-tokenizer --output-dir output/w2-runtime

- name: "integration:runtime-model"
id: runtime_model
timeout-minutes: 15
env:
OMP_NUM_THREADS: "1"
OPENBLAS_NUM_THREADS: "1"
MKL_NUM_THREADS: "1"
HF_HUB_OFFLINE: "1"
TRANSFORMERS_OFFLINE: "1"
run: |
"$QWEN3_PY" probes/w2_runtime_model/run.py \
--tokenizer-dir output/qwen3-tokenizer --output-dir output/w2-runtime-model

- name: "numeric:ir-qwen3-small"
id: qwen3_small
timeout-minutes: 15
Expand Down Expand Up @@ -210,6 +274,39 @@ jobs:
"$QWEN3_PY" probes/w2_qwen3_small/riscv.py \
--model-dir output/qwen3-small --output-dir output/qwen3-riscv

- name: Summarize W2 unit gates
if: always()
env:
FRONTEND_OUTCOME: ${{ steps.frontend_ops.outcome }}
BACKEND_OUTCOME: ${{ steps.backend_ops.outcome }}
RUNTIME_OUTCOME: ${{ steps.runtime.outcome }}
RUNTIME_MODEL_OUTCOME: ${{ steps.runtime_model.outcome }}
run: |
{
echo "## W2 frontend, backend and host runtime"
echo "unit:frontend-ops: $FRONTEND_OUTCOME"
echo "unit:backend-ops: $BACKEND_OUTCOME"
echo "unit:runtime: $RUNTIME_OUTCOME"
echo "integration:runtime-model: $RUNTIME_MODEL_OUTCOME"
for report in output/w2-backend-ops/report.md output/w2-runtime/report.md output/w2-runtime-model/report.md; do
if [ -f "$report" ]; then cat "$report"; fi
done
echo "These gates do not execute full-model IR/QEMU or text generation."
} >> "$GITHUB_STEP_SUMMARY"

- name: Upload W2 unit evidence
if: always()
uses: actions/upload-artifact@v4
with:
name: w2-unit-gates-output
path: |
output/w2-backend-ops/
output/w2-runtime/
output/w2-runtime-model/report.*
output/w2-frontend-ops.xml
output/w2-runtime-tests.xml
if-no-files-found: ignore

- name: Upload RISC-V tensor evidence
if: always()
uses: actions/upload-artifact@v4
Expand Down Expand Up @@ -307,7 +404,7 @@ jobs:
# This downloads 1.24 GB and executes the 2.38 GB FP32 model. It is not
# counted as a passing gate on PRs where it did not execute.
if: github.event_name != 'pull_request'
runs-on: ubuntu-latest
runs-on: ubuntu-24.04
timeout-minutes: 60
steps:
- uses: actions/checkout@v4
Expand Down Expand Up @@ -352,3 +449,65 @@ jobs:
name: qwen3-full-onnx-probe-output
path: output/qwen3-full-probe/
if-no-files-found: ignore


full-qwen3-frontend:
name: "W2 / complete Qwen3 frontend"
# Required on PRs as well as manual/nightly runs. No full-model ORT or IR inference.
runs-on: ubuntu-24.04
timeout-minutes: 30
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5
with:
python-version: "3.12"
- name: Install pinned frontend audit dependencies
run: |
python3 -m pip install "numpy==2.2.6" "onnx==1.18.0" "onnxruntime==1.22.1" "protobuf==5.29.5"
python3 -m pip check
- name: "frontend:parse-qwen3"
id: full_frontend
env:
OMP_NUM_THREADS: "1"
OPENBLAS_NUM_THREADS: "1"
run: |
set -euo pipefail
python3 "$PROBE_QWEN3_PARSE" --mode download \
--model-dir output/qwen3-full-model \
--output-dir output/qwen3-parse --timeout 1200
- name: Summarize complete Qwen3 frontend gate
if: always()
env:
PROBE_OUTCOME: ${{ steps.full_frontend.outcome }}
run: |
{
echo "## frontend:parse-qwen3"
echo "Outcome: $PROBE_OUTCOME"
if [ -f output/qwen3-parse/report.md ]; then
cat output/qwen3-parse/report.md
fi
echo "Complete graph and binding audit only; no full IR or QEMU numerical claim."
} >> "$GITHUB_STEP_SUMMARY"
- name: Upload complete Qwen3 frontend evidence
if: always()
uses: actions/upload-artifact@v4
with:
name: qwen3-full-frontend-parse-output
path: output/qwen3-parse/
if-no-files-found: ignore

w2-acceptance:
name: "W2 / required acceptance"
needs: [llm-deploy, full-qwen3-frontend]
if: always()
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- name: Require actual success from every W2 job
env:
W2_JOB_RESULTS: ${{ toJSON(needs) }}
run: python3 scripts/check_w2_ci_results.py >> "$GITHUB_STEP_SUMMARY"
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@ build/
models/
venv/
.venv/
.venv-linux/
.venv-*/
.claude/
.omo/
scratchv.egg-info/
Expand Down
5 changes: 5 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,11 @@ pytest tests/ -v # run all tests
- User-facing docs go in `docs/`.
- Inline code comments are for *why* not *what*.
- The README is the single source of truth for project-wide docs.
- Qwen3 deployment (W1 and later) is delivered and independently reproduced on
Linux. Use Ubuntu 24.04, Bash, Python 3.12 and Linux tool paths in public guides,
PR instructions and CI. See [the Linux reproduction contract](docs/llm-deploy-v1.0/LINUX_REPRODUCTION.md).
Personal development environments are not acceptance evidence; preserve the
actual platform and commit identity of historical measurements.

## Code of Conduct

Expand Down
32 changes: 22 additions & 10 deletions benchmarks/test_regalloc/regalloc.md
Original file line number Diff line number Diff line change
Expand Up @@ -284,20 +284,32 @@ convert_onnx_to_llvm(model) → LLVM IR (866K lines, 183MB)

## 7. 使用方式

以下复现命令面向 Linux / Bash,在仓库根目录执行。先准备 Python 3.12 虚拟环境和依赖:

```bash
python3.12 -m venv .venv-regalloc
source .venv-regalloc/bin/activate
python -m pip install -e ".[all]"
```

CNN 用例还需要 `models/graph/cnn.onnx`;请先确认模型文件存在。LLVM 对比依赖
`llvmlite`,报告中的 `llvm_available=false` 表示该项未执行,不能作为 LLVM 对比通过。

```bash
# 运行全部 4 项 benchmark,生成三格式报告(Windows PowerShell)
& .\.venv\Scripts\python.exe -m benchmarks.test_regalloc.bench_regalloc_linear `
--repeats 30 `
--output-json report.json `
--output-html report.html `
--output-md report.md
# 运行全部 4 项 benchmark,生成三格式报告
mkdir -p benchmark_reports
python -m benchmarks.test_regalloc.bench_regalloc_linear \
--repeats 30 \
--output-json benchmark_reports/regalloc_bench.json \
--output-html benchmark_reports/regalloc_bench.html \
--output-md benchmark_reports/regalloc_bench.md

# 单独运行某项
& .\.venv\Scripts\python.exe -m benchmarks.test_regalloc.bench_simple --repeats 100
& .\.venv\Scripts\python.exe -m benchmarks.test_regalloc.bench_dense --repeats 50
& .\.venv\Scripts\python.exe -m benchmarks.test_regalloc.bench_cnn `
python -m benchmarks.test_regalloc.bench_simple --repeats 100
python -m benchmarks.test_regalloc.bench_dense --repeats 50
python -m benchmarks.test_regalloc.bench_cnn \
--cnn-path models/graph/cnn.onnx --repeats 30
& .\.venv\Scripts\python.exe -m benchmarks.test_regalloc.bench_pseudo --repeats 30
python -m benchmarks.test_regalloc.bench_pseudo --repeats 30
```

---
Expand Down
Loading
Loading