The program

Every talk, one place

A curated archive of sessions delivered across our US editions — filter by track to find your next deep dive.

Keynote

The Lakehouse Decade: Where Storage Meets Intelligence

A decade-long look at how open table formats reshaped the modern data stack and what comes after the lakehouse.

Dr. Amara Okafor San Francisco
Jul 22, 2025 · 12:30 PM
AI / ML

Serving Trillion-Token Models on a Budget

Practical patterns for batching, quantization, and autoscaling LLM inference without burning your cloud bill.

Rohan Mehta Austin
Jul 23, 2025 · 3:00 PM
Data Engineering

Exactly-Once at a Million Events per Second

Lessons from running mission-critical streaming pipelines with strict ordering and replay guarantees.

Lin Zhao Chicago
Jul 17, 2024 · 3:00 PM
Cloud

FinOps for Multi-Cloud: Stop Guessing, Start Governing

A governance framework that gives finance and engineering one shared, real-time view of spend.

Erik Larsen Seattle
Jul 16, 2024 · 12:30 PM
AI / ML

Foundation Models as Infrastructure

Why the next platform layer is model-shaped, and how to build a reusable ML platform around it.

Dr. Valentina Cruz New York
Jul 18, 2023 · 9:00 AM
Analytics

Vectorized Query Engines from First Principles

A ground-up tour of columnar execution, SIMD, and the design tradeoffs behind modern OLAP engines.

Marcus Bell Boston
Jul 17, 2023 · 12:30 PM
Data Engineering

Data Contracts in Practice

How explicit producer–consumer contracts cut pipeline incidents by 60% across hundreds of teams.

Dr. Amara Okafor Denver
Jul 11, 2022 · 2:30 PM
AI / ML

Real-Time Feature Stores Without the Pain

Bridging the online/offline gap with a single source of truth for features and labels.

Lin Zhao Atlanta
Jul 12, 2022 · 10:30 AM
Keynote

The Warehouse Is the New Operating System

From dashboards to applications: how the cloud warehouse became a programmable platform.

Marcus Bell Las Vegas
Jul 20, 2021 · 10:30 AM
AI / ML

Building Feature Platforms That Scale to Thousands of Models

A blueprint for feature stores that stay fast, consistent, and debuggable as model count explodes.

Priya Nair San Francisco
Jul 22, 2025 · 9:00 AM
Data Engineering

Consensus at Planet Scale

What it really takes to run replicated, strongly-consistent storage across continents.

Kenji Tanaka San Francisco
Jul 21, 2025 · 11:00 AM
Analytics

Causal Inference for Product Decisions

Moving teams beyond correlation dashboards to causal models that actually inform decisions.

Sofia Russo Seattle
Jul 15, 2024 · 10:30 AM
Analytics

Query Optimization in the Age of Open Formats

How cost-based optimizers are being rebuilt for Iceberg, Delta, and the disaggregated stack.

David Goldberg Seattle
Jul 15, 2024 · 2:30 PM
AI / ML

Evaluating LLMs You Can Actually Trust

A practical framework for building evaluation suites that catch real regressions, not vibes.

Dr. Aisha Rahman Seattle
Jul 17, 2024 · 9:00 AM
Cloud

Golden Paths: Platform Engineering on Kubernetes

Designing internal developer platforms that abstract Kubernetes without hiding what matters.

Carlos Mendez New York
Jul 19, 2023 · 9:00 AM
AI / ML

Real-Time Recommenders at Consumer Scale

Serving large embedding models and fresh features with single-digit-millisecond latency.

Mei Chen New York
Jul 18, 2023 · 2:30 PM
Data Engineering

Observability for Data Pipelines

Lineage, freshness, and anomaly detection — the telemetry that keeps data products reliable.

James O'Brien Atlanta
Jul 13, 2022 · 10:30 AM
AI / ML

Federated Learning Without the Hype

Where privacy-preserving training genuinely pays off — and where a central warehouse still wins.

Dr. Fatima Al-Sayed Atlanta
Jul 11, 2022 · 4:30 PM
Data Engineering

Data Quality as a First-Class Product

Treating freshness, completeness, and contracts as SLAs your consumers can depend on.

Wei Liu Las Vegas
Jul 19, 2021 · 1:00 PM
Keynote

The Road to the Lakehouse

A look back at how warehouses and lakes converged, and why open formats won the decade.

Marcus Bell Virtual
Jul 15, 2020 · 1:00 PM
AI / ML

Scaling ML Platforms When Everyone Went Remote

Lessons from operating shared ML infrastructure for distributed teams overnight.

Priya Nair Virtual
Jul 16, 2020 · 12:30 PM
Data Engineering

Schema Evolution in Open Table Formats

Safe, backward-compatible schema change strategies for petabyte-scale tables.

Dr. Amara Okafor Chicago
Jul 16, 2019 · 12:30 PM
Cloud

Elastic Compute for Bursty Analytics

Autoscaling analytical workloads without paying for idle capacity between the bursts.

Erik Larsen Las Vegas
Jul 12, 2018 · 11:00 AM
Keynote

Foundations of the Cloud Data Warehouse

The opening keynote that launched the conference: separating storage from compute, for good.

Dr. Valentina Cruz San Jose
Sep 19, 2017 · 1:00 PM
Keynote

Real-Time Sync: Event-Driven Integration Across Enterprise Systems

Replacing brittle scheduled ETL with change-data-capture and streaming so account updates propagate to every downstream system in seconds, not hours.

Srikanth Jonnakuti San Francisco
Jul 23, 2025 · 9:00 AM
Data Engineering

Self-Serve Data Platforms That Don't Break Trust

How governed, self-serve platforms let analysts ship fast while keeping lineage and access under control.

Naomi Clarke San Francisco
Jul 22, 2025 · 11:00 AM
AI / ML

Tail Latency Is the Product: Serving LLMs in Real Time

Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.

Arjun Patel San Francisco
Jul 22, 2025 · 11:00 AM
Analytics

Experimentation at Scale Without Fooling Yourself

Designing A/B systems that account for interference, peeking, and the messy reality of product metrics.

Hana Kim San Francisco
Jul 21, 2025 · 1:00 PM
AI / ML

Embeddings in Production: Retrieval That Stays Fresh

Keeping vector indexes consistent and low-latency as embeddings and documents change constantly.

Lucia Fernandez San Francisco
Jul 21, 2025 · 2:30 PM
Keynote

Reading the Cloud Bill: FinOps for Engineers

Turning raw billing exports into engineering decisions developers can actually act on.

Ben Fisher San Francisco
Jul 21, 2025 · 2:30 PM
AI / ML

Golden Paths for Production ML

Making the common ML workflows boring: paved roads that turn heroics into routine releases.

Anita Desai San Francisco
Jul 21, 2025 · 9:00 AM
Data Engineering

Change Data Capture at High Volume

Keeping operational databases and the lakehouse in sync without melting the source system.

Omar Haddad Seattle
Jul 16, 2024 · 10:30 AM
Data Engineering

On-Call for Data: Building a Reliability Culture

The tooling, runbooks, and incentives that keep data products dependable and engineers rested.

Terrence Walker Seattle
Jul 16, 2024 · 3:00 PM
AI / ML

Benchmarks You Can Trust

Why most ML benchmarks mislead, and how to build evaluation suites that survive contact with reality.

Dr. Katherine Wolfe Seattle
Jul 17, 2024 · 4:30 PM
Data Engineering

Storage Engines for Geo-Distributed Databases

Design tradeoffs in LSM trees, replication, and consensus for databases that span continents.

Daniel Park Seattle
Jul 17, 2024 · 9:00 AM
Cloud

Secure Multi-Cloud Networking for Data Teams

Connecting data platforms across providers without opening holes your security team will hate.

Diego Ramirez Seattle
Jul 17, 2024 · 10:30 AM
AI / ML

Fairness Guardrails That Ship

Translating fairness research into measurable guardrails teams can actually put into production.

Dr. Grace Adeyemi New York
Jul 17, 2023 · 3:00 PM
Data Engineering

The Modern Ingestion Stack

Batch, streaming, and CDC under one roof — choosing the right tool for each source without sprawl.

Naomi Clarke New York
Jul 19, 2023 · 1:00 PM
Analytics

Causal Measurement for Growth Teams

Geo experiments, switchbacks, and synthetic controls when a clean A/B test just isn't possible.

Hana Kim New York
Jul 18, 2023 · 12:30 PM
AI / ML

Scaling the ML Platform Team

Org design, interfaces, and the platform-as-product mindset behind ML infrastructure that scales.

Anita Desai New York
Jul 17, 2023 · 4:30 PM
Data Engineering

Pipeline Reliability SLAs in Practice

Defining freshness and completeness SLAs your consumers can plan around — and meeting them.

Terrence Walker Atlanta
Jul 13, 2022 · 2:30 PM
Data Engineering

Streaming Ingestion Without the 3 a.m. Pages

Backpressure, schema drift, and dead-letter handling for ingestion pipelines that stay calm.

Omar Haddad Atlanta
Jul 11, 2022 · 1:00 PM
AI / ML

Vector Search Before It Was Cool

Hard-won lessons building approximate nearest-neighbor systems for recommendations at scale.

Lucia Fernandez Atlanta
Jul 11, 2022 · 10:30 AM
Cloud

Cost Optimization for the Lakehouse

Right-sizing compute, tuning file layouts, and killing the silent spend in your data platform.

Ben Fisher Atlanta
Jul 13, 2022 · 10:30 AM
AI / ML

Evaluating Models When Everyone Works Remote

Reproducible, shareable evaluation pipelines for teams that no longer sit in the same room.

Dr. Katherine Wolfe Virtual
Jul 20, 2021 · 2:30 PM
Data Engineering

Consensus Protocols, Explained for Builders

A practical tour of Raft, Paxos variants, and the failure modes that bite real deployments.

Daniel Park Virtual
Jul 19, 2021 · 2:30 PM
Cloud

Multi-Cloud by Necessity, Not Hype

When spanning providers genuinely pays off — and the networking realities nobody warns you about.

Diego Ramirez Virtual
Jul 15, 2020 · 1:00 PM
Data Engineering

Building Governed Platforms Overnight

Lessons from standing up governed, remote-first data platforms when the office disappeared.

Naomi Clarke Virtual
Jul 16, 2020 · 3:00 PM
AI / ML

Inference Economics in the Early GPU Crunch

Batching and model selection strategies from the days when every GPU hour really hurt.

Arjun Patel Chicago
Jul 18, 2019 · 12:30 PM
Data Engineering

Reliability Lessons from a Decade of Outages

What recurring incidents taught one team about lineage, alerting, and blameless postmortems.

Terrence Walker Chicago
Jul 18, 2019 · 10:30 AM
Cloud

Elastic Storage for the Bursty Enterprise

Decoupling storage growth from compute cost when workloads are spiky and unpredictable.

Ben Fisher Las Vegas
Jul 10, 2018 · 1:00 PM
Analytics

Early Experimentation Platforms

Building the first in-house A/B platform: metrics, guardrails, and the trust it took to earn.

Hana Kim Las Vegas
Jul 11, 2018 · 10:30 AM
AI / ML

Measuring Models in the Warehouse Era

An early call for rigorous, reproducible model evaluation as data moved into the cloud warehouse.

Dr. Grace Adeyemi San Jose
Sep 19, 2017 · 11:00 AM
Data Engineering

Designing Storage for the New Warehouse

The storage-engine decisions behind separating compute from durable, columnar data.

Daniel Park San Jose
Sep 21, 2017 · 11:00 AM
Cloud

Intelligent Automation for Cloud-Native Integration Platforms

How model-assisted automation reduced manual reconciliation and incident toil across a multi-account, event-driven integration estate.

Srikanth Jonnakuti San Francisco
Jul 22, 2025 · 3:00 PM
Cloud

Secure DevSecOps for Multi-Account Cloud Infrastructure

Policy-as-code, automated guardrails, and infrastructure templates that keep enterprise cloud estates compliant without slowing delivery.

Diego Ramirez San Francisco
Jul 21, 2025 · 4:30 PM