New models
Highlights
PD instances can switch between prefill and decode on the fly, no restart needed.
The prefix cache now runs on a Rust core by default.
DeepSeek-V4.1 gets 22% faster first token on long prompts.
Kimi K3 gets 20.6% higher prefill throughput in PD serving.







