Skip to content

Milestones

List view

  • 0.2.0: speculative decoding grows from a shipped feature into a subsystem (rejection sampling #512, shared primitives #443), GLM5.2 decode serving converges on the vLLM DP8/EP8 reference (#542 + perf backlog, #590 DSpark×prefix-cache), batch-invariant Qwen3 inference reaches its e2e bar (#435), and Prometheus observability stabilizes on the Qwen3 line (#602: real counters + Grafana dashboard). Plus frontend error-surface polish (#294, #584).

    No due date
    5/19 issues closed