Skip to content
ust-xuPublic

About

[ICLR'26] Official code for TD-MoE: tensor decomposition for compressing Mixture-of-Experts language models.

Resources

Stars

13 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

TD-MoE: Tensor Decomposition for MoE Models

TD-MoE cover

ICLR 2026 Python 3.10+ PyTorch 2.0+ License

Paper

This repository provides the official implementation of TD-MoE, a data-aware tensor decomposition framework for compressing Mixture-of-Experts language models by jointly decomposing experts within each layer.

Installation

git clone https://github.com/ust-xu/TD-MoE.git
cd TD-MoE

pip install -r requirements.txt

cd lm-evaluation-harness
pip install -e .
cd ..

Quick Start

Set model.path in the config file or override it from the command line, then run the full TD-MoE pipeline:

# Qwen2-57B-A14B with 20% compression
python run.py --config configs/qwen2/qwen2_0.2.yaml --model_path /path/to/Qwen2-57B-A14B

# Mixtral-8x7B-v0.1 with 20% compression
python run.py --config configs/mixtral/mixtral_0.2.yaml --model_path mistralai/Mixtral-8x7B-v0.1

Pipeline

TD-MoE executes a staged compression pipeline controlled by each YAML config:

  1. covariance: collect activation statistics from calibration data.
  2. whitening: compute whitening transforms and decompose selected MoE layers.
  3. evaluate: run perplexity and downstream benchmark evaluation.

You can control the executed stages through the pipeline.steps field in each config.

Configuration Notes

  • compression.ratio: target parameter reduction ratio.
  • compression.whiten_type: whitening mode used during decomposition.
  • compression.whitening_nsamples: number of calibration samples.
  • evaluation.eval_tasks: downstream evaluation tasks passed to the harness.
  • output.save_path: directory for intermediate artifacts and evaluation results.

Citation

@inproceedings{tdmoe2026,
  title={TD-MoE: Tensor Decomposition for MoE Models},
  author={Xu, Yuebin and Wang, Yanhong and Peng, Xuemei and Zang, Hui and Chen, Minghao and Xia, Pengfei and Wen, Zeyi},
  booktitle={International Conference on Learning Representations (ICLR)},
  year={2026},
  url={https://openreview.net/pdf?id=D9cnZNZfxX}
}

Acknowledgments

This repository builds on several excellent open-source projects:

About

[ICLR'26] Official code for TD-MoE: tensor decomposition for compressing Mixture-of-Experts language models.

Resources

Stars

13 stars

Watchers

0 watching

Forks

Contributors

Languages