Tail Latency Is the Product: Serving LLMs in Real Time
Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.
San Francisco
Jul 22, 2025 · 11:00 AM

Speaker
Senior Software Engineer · Northwind AI
Arjun builds high-throughput LLM serving systems at Northwind AI, obsessing over tail latency, batching, and squeezing every token out of each GPU.
Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.
San Francisco
Jul 22, 2025 · 11:00 AM
Batching and model selection strategies from the days when every GPU hour really hurt.
Chicago
Jul 18, 2019 · 12:30 PM
A. Patel, M. Bell
AI Systems · 2025 · 12 citations