Serving Trillion-Token Models on a Budget
Practical patterns for batching, quantization, and autoscaling LLM inference without burning your cloud bill.
Austin
Jul 23, 2025 · 3:00 PM

Speaker
Principal Data Scientist · Northwind AI
Rohan builds large-scale LLM serving infrastructure at Northwind AI. His talks on cost-aware inference have become required viewing for teams shipping generative AI to production.
Practical patterns for batching, quantization, and autoscaling LLM inference without burning your cloud bill.
Austin
Jul 23, 2025 · 3:00 PM
R. Mehta, V. Cruz
AI Systems · 2025 · 18 citations
R. Mehta, V. Cruz
AI Systems · 2025 · 19 citations
R. Mehta, X. Dubois
AI Systems · 2025 · 8 citations
D. Chen, R. Mehta
Data Engineering · 2025 · 4 citations
L. Zhao, V. Cruz, R. Mehta
ML Platforms · 2022 · 118 citations
R. Mehta, K. Nakamura
DevOps · 2021 · 82 citations
R. Mehta, I. Santos
Cloud · 2019 · 183 citations