New models
Highlights
Built for large-scale deployments, delivering reliable, low-latency, high-throughput serving from a single GPU to distributed clusters.
Supports a wide range of open models — from LLMs to diffusion models — and runs across diverse hardware platforms.
Incorporates disaggregated prefill/decode, speculative decoding, parallelisms, a zero-overhead scheduler, and optimized GPU kernels.
Select your preferences and run the deployment command. SGLang is designed to be easy to install and deploy.
The easiest ways to get started.
Start the server with a single command pointing to your model.
Use standard OpenAI-compatible endpoints to interact with your model.
uv pip install sglang sglang-kernel \
--extra-index-url https://sgl-project.github.io/whl/cu129/ \
--extra-index-url https://download.pytorch.org/whl/cu129 \
--index-strategy unsafe-best-matchA single engine that runs across various models and hardware.
New models
Highlights
New models
Highlights
With return_sampling_mask, each decode step returns the exact token support the sampler drew from and the log-probability of the sampled token under it, so a trainer can replay the rollout without reconstructing top-k or top-p. Masks now run under overlap…
Branching-point caching for the SWA component keeps the sliding-window state at the point where requests fork from a shared prefix, so branches reuse it instead of recomputing. On DeepSeek-V4-Flash with a shared system prompt, token hit rate rises from 43.8%…
A DCP1 prefill can now transfer its DSpark draft KV to a DCP-N decode, so hybrid models such as Kimi-Linear run DSpark in disaggregated, context-parallel serving. Verified on 8x B300 over NIXL and Mooncake up to 256K input.
/v1/responses no longer retains results in memory unless the server starts with --enable-response-store. Without it, retrieval, previous_response_id chaining, and background requests return 400; PD deployments cannot enable it.
A CPU-only simulator runs the real scheduler, radix cache, and hierarchical cache with a latency predictor in place of the model forward. Against measured serving traces it predicts TTFT within about 6% on most traces (up to 10% on the longest 32K to 128K…
The strategy-based implementation is now the only prefill CP path; the v1 runtime and its CLI options are gone. Prefill CP on HIP, NPU, and MUSA is rejected until those platforms are ported.
New models
Highlights
SGLang can now do beam search. Pass beam_width in your request and you get back the n best sequences instead of a single sample. It works out of the box next to regular requests, though it does not yet mix with speculative decoding, disaggregation, DP…
DeepEP's new ElasticBuffer engine is available as --moe-a2a-backend deepep_v2 for DeepSeek-V3/V4 and Qwen3-MoE in FP8. Its buffers have a fixed size, so decode can run under CUDA graphs even across nodes. Performance is on par with the classic backend.
With --enable-layernorm-sp, each tensor-parallel rank normalizes only its own share of the prefill tokens instead of all of them. That takes 3.5% off Qwen3-8B prefill on H100 and 5.6% on B200, and the saving grows with the TP degree. Dense Qwen3 models only…
If you serve MXFP4 experts on Hopper, you can now quantize the activations to FP8 as well with --flashinfer-mxfp4-moe-precision fp8. DeepSeek-V4-Flash gains about 12% output throughput with no change in GSM8K accuracy. Needs FlashInfer 0.6.18.
Decode context parallelism now runs on trtllm_mla, not just CuTe DSL and Tokenspeed. It pays off at long context: at 128K input, plain TP stops scaling around 680 tokens per second on eight B200s, while DCP keeps going as concurrency grows.
DSA prefill top-k moves to the v2 kernel, 1.3 to 1.8 times faster on B200. KDA models get an opt-in fused accept path, SGLANG_OPT_KDA_FUSED_ACCEPT_STATE=1, that cuts MTP verify-and-commit time by 45% to 63% on Kimi-Linear shapes with bit-identical output.
New models
Highlights
Checkpoint pages now stage from storage while CUDA graphs capture. Qwen3-32B on H100 starts 8.6-11.7% faster than serial with prefetch, and 2.38x faster (35.6s vs 84.8s) than the plain default. Opt in with --startup-weight-load-mode overlap.
The TP LMHead's allgather + scatter becomes a single all-to-all for pure-DP dp-attention. On DeepSeek-V4-Pro B200 decode, LMHead time drops 320us to 169us and TPOT improves 36.97ms to 35.67ms.
Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL. DeepSeek-V4-Flash TP4 decode on Blackwell gains up to +6.9% at small batches. Auto-enabled for DeepSeek-V3/V3.2/V4; elsewhere…
--quantization quark_mxfp4 dequantizes ModelOpt and Quark NVFP4 weights and requantizes to MXFP4 at load, never holding a full-precision copy. 97.5-100.2% GSM8K recovery vs the NVFP4 reference across MiniMax-M2.7, GLM-5.1, Kimi-K2.6, Qwen3.5-397B, and…
A grouped-head MLA verify kernel replaces the MHA-shaped split-KV path that re-read the shared latent once per head, for 1.37-1.77x throughput and 1.45-2.42x ITL at concurrency 2-32; the AITER MLA prefill kernel now accepts K3's 12-head shape (TTFT up to…
Triton, FlashInfer, Inductor, DeepGEMM, and CUDA driver caches all move under SGLANG_CACHE_DIR. The first launch after upgrading recompiles once; see Breaking Changes.
Highlights
A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint. SGLang…
MiniMax's video generation model that produces a video and a synchronized stereo audio track in one request, served natively on SGLang-Diffusion across all three public task profiles: text-to-video-and-audio (t2va), first/last-frame conditioning (fl2va)…
Migrates the front half of the server, everything from network ingress up to the point a tokenized request is handed to the GPU scheduler, from Python to a multi-threaded Rust implementation.
The DeepSeek-MLA decode context-parallel path gains pluggable comm backends. a2a exchanges packed attention output plus fp32 LSE in a single NCCL collective per layer, with fp8 KV carried as uint8 byte transport; fi_a2a delegates the cross-rank exchange…
A new prefill parallelism strategy that prefetches peer expert weights over NVLink P2P and computes all experts locally, removing EP all-to-all token dispatch. On 4x B200 with gpt-oss-120b, prefill-only, DWDP4 reaches 1.92x over DEP4 at MNT 32K / ISL 32K, and…
For agentic and RL-rollout workloads, requests can carry a stable session_id so eviction knows which prefixes an active session still references, instead of evicting purely by cache policy. Release the references with /close_session. Opt in with…
Highlights
A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length. Reaches 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro, TP8 on B300 (bs=1). Enable with…
A 975B-parameter multimodal MoE with a 1M-token context. It mixes sliding-window, full and Mamba2 linear attention, and adds an NVFP4 MoE, optional vision/audio towers and native MTP. On Blackwell it reaches up to 71.7k tok/s input and 171.0 tok/s per-user…
UnifiedRadixTree is now the default for SWA, Mamba and DSA models. Replay SSM and Mamba int8 checkpoints are synced onto it, and a cache hit now resets only the state it used.
KV and indexer cache layers are sharded across CP ranks. Each rank owns a disjoint layer range instead of all layers. That cuts per-rank KV memory by ~74% (0.77 to 0.20 GB/rank) at 8192 tokens on GLM-5.2-FP8, 78 layers, cp_size=4. Enable with…
Drops the per-draft SSM snapshot. Speculative scratch goes from 11.5 GB to 1.8 GB per GPU (6.4x smaller) on Qwen3.5-35B-A3B at TP1, at accuracy and throughput parity. Opt in with --enable-gdn-replayssm-spec (default off; GDN with a linear draft chain only…
The first correct KDA MTP path. Its recurrent_kda decode kernel runs at 29.6 us vs 36.8 us for Triton (ncu, B=64). The full decode path reaches parity by B=128 and 1.35x at B=256, and is slower below that. Separately, GDN/KDA CuteDSL prefill fuses state I/O…
43 new contributors
New models
Highlights
We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving. It now runs at 500+ tok/s/user on 8x B300, 450 on 4x GB300 (bs=1). Run GLM-5.2 with our.
Built-in web_search support backed by Exa.
Breakable CUDA Graph is now the default capture path, reducing per-step kernel-launch overhead; full CUDA Graph support for the prefill phase lands as experimental.
New FlashKDA prefill backend for safe-gate KDA linear attention, plus ReplaySSM buffered output-only decode for linear attention.
Adds FlashInfer all-to-all with the flashinfer_trtllm_routed MoE runner.
Optimizes C128 state-pool allocation using the request state pool; FlashMLA sparse prefill is now enabled by default for DeepSeek-V4, reaching >10% throghput gain on long context; Non paged indexer support for long context prefill, with >5% e2e throughput gain
55 new contributors
New models
Highlights
5x higher throughput at the same interactivity, serving DeepSeek-V4 on NVIDIA GB300 with SGLang.
Two dispatch-time load-balancing methods for DeepEP expert parallelism: Waterfill for shared-expert dispatch and LPLB for redundant expert replicas, improving throughput for DeepSeek-V3/R1 and DeepSeek-V4.
New CuteDSL prefill kernel for Kimi-Linear (KDA), 1.08-1.52x faster than the Triton path via a reusable scratch workspace, plus a cuda-graph padding fix; see the Kimi-Linear cookbook.
An int8 checkpoint pool stores recurrent states compactly in the Mamba radix cache, substantially increasing prefix-cache capacity for KDA / GDN models; the speculative conv-window intermediate cache is deduplicated with a sliding-window layout, halving its…
Balances token routing across redundant expert replicas by solving a per-layer LP; opt-in via --ep-dispatch-algorithm=lp, default behavior unchanged.
MSCCL++ migrates to the upstream mscclpp Python package (Executor + DSL compiler) with auto-tuned collectives for TP=8 single-node and TP=16 two-node; FlashInfer fused allreduce + residual + RMSNorm re-enables an MNNVL backend behind…
70 new contributors
New models
Highlights
Tree drafting with topk > 1 is production-ready across the triton / FA3 / MLA / aiter backends, including page_size > 1 and Mamba/hybrid-linear models. Spec V1 is deprecated, with EAGLE/MTP now running on the unified V2 worker, and topk = 1 drafting is…
Unified async value passing through FutureMap plus moving prefill input transfer onto the forward stream reduced per-step launch overhead and improved stability under high concurrency.
Piecewise (PCG) and Breakable (BCG) CUDA Graph capture more of the model to cut per-step kernel-launch overhead, now extended to DSA models, Kimi-K2.5, and DeepSeek V4.
New FlashInfer Gated DeltaNet (GDN) kernels and a CuTeDSL GDN prefill kernel speed up Qwen 3.5 on Blackwell GPUs.
HybridModel (SWA/Mamba) launches HiCache through UnifiedTree by default, bringing hierarchical KV-cache offload to sliding-window and Mamba hybrids out of the box.
Offload VLM vision encoding onto Intel Xeon CPUs alongside GPUs, with up to ~1.3x P99 TTFT and request-throughput gains under load.
35 new contributors
New models
Highlights
Full inference path for DeepSeek-V4, including: Parallelism: Tensor Parallelism/Expert Parallelism/Context Parallelism/Data Parallel Attention; Hardware: Nvidia B300/B200/H200/H100/GB200/GB300, AMD MI35X; Prefill-Decode Disaggregation; HiSparse for offloading…
New MLA prefill/decode kernels integrated as an attention backend on SM100, with FP8 KV cache support for low-latency MLA serving
PDL enabled across DSv3.2 / GLM-5 kernels, torch.mm for the DeepSeek V3.2 indexer GEMM, and a reland of the Cute-DSL FP4 dense GEMM — materially trimming low-latency overheads on FP4 paths
HiCache framework support for UnifiedRadixTree (with SWA), HiCache for DeepSeek V4, SSD offload through Mooncake store, and stability fixes across cascade eviction, tombstone replay, and partial-match paths
Adaptive Spec V2, EAGLE-3 SWA + newer drafters, Kimi K2.5 EAGLE-3 MLA, Gemma 3/4 + EAGLE-3, and an extensive naming / shape-handling refactor across draft-extend paths
Gateway DeepEP source swapped from a community fork to deepseek-ai/DeepEP@hybrid-ep so DeepEP builds and runs cleanly on the CUDA 13 default; FlashInfer pinned at 0.6.11.post1 alongside a gpt-oss triton-kernel fix
71 new contributors
New models
Highlights
Default CUDA version moves to 13.0 across SGLang, sgl-kernel, and Docker images, and PyTorch is upgraded from 2.9 to 2.11 — modernizing the build matrix and unlocking newer kernels (tracking issue)
Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths
Decode-side prefix caching now works under prefill/decode disaggregation, recovering radix-cache hit rates and TTFT savings for long shared prefixes in disaggregated deployments
New high-throughput spec-decode kernel from the kernel community, expanded across model backends and AMD ROCm
Drop-in FA3 kernels contributed by the community, integrated alongside FA4 to give users a high-performance option that's easy to maintain
LoRA now works on the largest MLA-based MoE models, including DeepSeek-V3 MLA LoRA and Kimi K2 — enabling adapter-based fine-tuning of frontier-scale models
77 new contributors
New models
Highlights
Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns
Integrate Elastic NIXL-EP into SGLang, enabling partial failure tolerance for DeepSeek MoE deployments — when a GPU fails, the system redistributes expert weights and continues serving without full restart
Gathers scattered head slices into contiguous memory for bulk RDMA transfer, reducing RDMA request count on GQA models by ~1000x. TPS/GPU on large concurrency increased by ~5x with Prefill TP4+Decode DEP4 on Qwen3.5
Integrate HiSparse sparse attention backend for efficient long-context inference with reduced compute through sparsity-aware attention
Model support: LTX-2, Hunyuan3D-2, Helios; Performance improvements on Qwen-image, Z-image increased by 1.5x; New platform: macOS; New feature: enhance the performance of diffusers backend by integrating all optimization from Cache-DiT; SKILLs: feel free to…
Integrate FlashInfer mxfp8 kernels for GEMM and MoE operations, enabling mixed-precision FP8 inference with higher accuracy through microscaling for RL and general workloads
From first-time users to teams debugging complex deployments, the community is open to everyone.







