Hannah works on inference infrastructure at Lattice & Loom. She spends her days making GPUs do more with less and her evenings explaining the bill.
SESSIONS · 2
15:00–15:45
TUE · DEC 1
MAIN STAGE
AI in Production: Beyond the Demo
AI & ML·Panel (45 min)·with Aisha Bello, Tomas Lindqvist
Everyone has an AI demo; far fewer have an AI feature that survives real users. This panel brings together three practitioners shipping LLM-powered products to talk honestly about evaluation, cost, guardrails, and what changed in their engineering practice. Moderated, with plenty of time for audience questions.
10:45–11:30
WED · DEC 2
MAIN STAGE
Inside the Token Budget: Serving LLMs Without Burning Money
AI & ML·Deep Dive (45 min)
Our inference bill grew 40x in a year — then we cut it by two thirds without hurting quality. A deep dive into the mechanics of LLM serving costs: caching and batching strategies, model routing, quantization trade-offs, and the dashboards that actually predict spend. Real numbers, real regrets.