This repository provides the official implementation of TD-MoE, a data-aware tensor decomposition framework for compressing Mixture-of-Experts language models by jointly decomposing experts within each layer.
git clone https://github.com/ust-xu/TD-MoE.git
cd TD-MoE
pip install -r requirements.txt
cd lm-evaluation-harness
pip install -e .
cd ..Set model.path in the config file or override it from the command line, then run the full TD-MoE pipeline:
# Qwen2-57B-A14B with 20% compression
python run.py --config configs/qwen2/qwen2_0.2.yaml --model_path /path/to/Qwen2-57B-A14B
# Mixtral-8x7B-v0.1 with 20% compression
python run.py --config configs/mixtral/mixtral_0.2.yaml --model_path mistralai/Mixtral-8x7B-v0.1TD-MoE executes a staged compression pipeline controlled by each YAML config:
covariance: collect activation statistics from calibration data.whitening: compute whitening transforms and decompose selected MoE layers.evaluate: run perplexity and downstream benchmark evaluation.
You can control the executed stages through the pipeline.steps field in each config.
compression.ratio: target parameter reduction ratio.compression.whiten_type: whitening mode used during decomposition.compression.whitening_nsamples: number of calibration samples.evaluation.eval_tasks: downstream evaluation tasks passed to the harness.output.save_path: directory for intermediate artifacts and evaluation results.
@inproceedings{tdmoe2026,
title={TD-MoE: Tensor Decomposition for MoE Models},
author={Xu, Yuebin and Wang, Yanhong and Peng, Xuemei and Zang, Hui and Chen, Minghao and Xia, Pengfei and Wen, Zeyi},
booktitle={International Conference on Learning Representations (ICLR)},
year={2026},
url={https://openreview.net/pdf?id=D9cnZNZfxX}
}This repository builds on several excellent open-source projects:
