All speakers
Arjun Patel

Speaker

Arjun Patel

Senior Software Engineer · Northwind AI

LLM ServingInference

Arjun builds high-throughput LLM serving systems at Northwind AI, obsessing over tail latency, batching, and squeezing every token out of each GPU.

Talks

AI / ML

Tail Latency Is the Product: Serving LLMs in Real Time

Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.

San Francisco

Jul 22, 2025 · 11:00 AM

AI / ML

Inference Economics in the Early GPU Crunch

Batching and model selection strategies from the days when every GPU hour really hurt.

Chicago

Jul 18, 2019 · 12:30 PM

Published papers

Contextual Embedding Caching for Low-Latency RAG Inference

A. Patel, M. Bell

AI Systems · 2025 · 12 citations