| | Pushing the Limits of Serving DeepSeek-V4-Pro (lmsys.org) |
| 1 point by gmays 8 days ago | past | discuss |
|
| | Miles v0.1: Production-level Post-training (lmsys.org) |
| 3 points by gmays 10 days ago | past | discuss |
|
| | Pushing the Limits of Serving DeepSeek-V4-Pro (lmsys.org) |
| 2 points by vzhou842 16 days ago | past |
|
| | Miles v0.1: Production-level Post-training (lmsys.org) |
| 2 points by nblintao 23 days ago | past |
|
| | The next generation of speculative decoding: DFlash and Spec V2 (lmsys.org) |
| 1 point by ronfriedhaber 42 days ago | past |
|
| | Agent-Assisted SGLang Development: An Initial Exploration (lmsys.org) |
| 1 point by gmays 66 days ago | past |
|
| | The next generation of speculative decoding: DFlash and Spec V2 (lmsys.org) |
| 4 points by gmays 83 days ago | past |
|
| | No Token Left Behind: Demystifying Token-in-Token-Out in Miles (lmsys.org) |
| 2 points by kkm 3 months ago | past |
|
| | DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles (lmsys.org) |
| 80 points by mji 4 months ago | past | 10 comments |
|
| | Pipeline Parallelism in SGLang: Scaling to Million-Token Contexts (lmsys.org) |
| 3 points by roody_wurlitzer 6 months ago | past |
|
| | Pipeline Parallelism in SGLang: Scaling to Million-Token Contexts and Beyond (lmsys.org) |
| 1 point by gmays 7 months ago | past |
|
| | Production-Ready Speculative Decoding Models and Framework (lmsys.org) |
| 1 point by gmays 8 months ago | past |
|
| | Mini-SGLang: Efficient Inference Engine in a Nutshell (lmsys.org) |
| 2 points by matt_d 8 months ago | past |
|
| | Power Up FSDP2 as a Flexible Training Back End for Miles (lmsys.org) |
| 1 point by gmays 9 months ago | past |
|
| | NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference (lmsys.org) |
| 115 points by yvbbrjdr 11 months ago | past | 93 comments |
|
| | Deploying DeepSeek on 96 H100 GPUs (lmsys.org) |
| 285 points by GabrielBianconi on Aug 29, 2025 | past | 80 comments |
|
| | Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP: 2.7x Throughput (lmsys.org) |
| 1 point by gmays on June 17, 2025 | past |
|
| | Match DeepSeek's inference system performance with SGLang (lmsys.org) |
| 1 point by echaozh on May 6, 2025 | past |
|
| | Does style matter? Disentangling style and substance in Chatbot Arena (lmsys.org) |
| 2 points by ZeljkoS on Feb 9, 2025 | past |
|
| | Faster JSON Decoding for LLMs (lmsys.org) |
| 1 point by gaocegege on Dec 18, 2024 | past |
|
| | Does style matter? Disentangling style and substance in Chatbot Arena (lmsys.org) |
| 1 point by scottfr on Aug 29, 2024 | past |
|
| | LLM Lookahead Decoding (lmsys.org) |
| 2 points by mr-ai on Aug 20, 2024 | past |
|
| | From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline – Lmsys Org (lmsys.org) |
| 2 points by swyx on Aug 3, 2024 | past |
|
| | Faster Open-Source Llama3 Serving with SGLang Runtime (vs. TensorRT-LLM, VLLM) (lmsys.org) |
| 4 points by yvbbrjdr on July 25, 2024 | past |
|
| | RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing (lmsys.org) |
| 4 points by adr1an on July 5, 2024 | past |
|
| | RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing (lmsys.org) |
| 4 points by not-chatgpt on July 1, 2024 | past | 1 comment |
|
| | Introducing Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by JumpCrisscross on June 21, 2024 | past |
|
| | Introducing Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by CharlesW on May 20, 2024 | past |
|
| | Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by imjonse on May 17, 2024 | past |
|
| | Whats up with Llama-3? LMSYS leaderboard analysis (lmsys.org) |
| 2 points by aadillpickle on May 15, 2024 | past |
|
|
| More |