Add Kimi K3 - #1626
Conversation
|
Independent validation on 4× M3 Ultra in TP4, plus one note for anyone applying this to the current release. We ran this PR's model file on a 4-node Apple silicon cluster and it works. Some What was actually tested — please read before weighing thisWe did not run the complete PR on So these fixtures validate our backport plus this PR's model and tool-parser Setup
What workedA four-fixture suite over the HTTP API, all passing:
Also: load ~5 min at 420.8 GB/rank; prefill throughput flat at ~108–114 tok/s Durability boundary: this validates model loading, TP4 sharding and
|
|
Thanks again for the Kimi-K3 implementation. We have a small tests-led follow-up The coverage is deliberately independent of the full checkpoint — no downloads,
We ran negative controls for wrong Tolerances are The three commits are independent and touch disjoint files, so they can be taken The PR already has good parser-level Kimi tool-call cases, so we have not |
Model: https://huggingface.co/moonshotai/Kimi-K3
Tested: https://huggingface.co/kernelpool/Kimi-K3-2bit-UVMAX