The Lakehouse Decade: Where Storage Meets Intelligence
A decade-long look at how open table formats reshaped the modern data stack and what comes after the lakehouse.
The program
A curated archive of sessions delivered across our US editions — filter by track to find your next deep dive.
A decade-long look at how open table formats reshaped the modern data stack and what comes after the lakehouse.
Practical patterns for batching, quantization, and autoscaling LLM inference without burning your cloud bill.
Lessons from running mission-critical streaming pipelines with strict ordering and replay guarantees.
A governance framework that gives finance and engineering one shared, real-time view of spend.
Why the next platform layer is model-shaped, and how to build a reusable ML platform around it.
A ground-up tour of columnar execution, SIMD, and the design tradeoffs behind modern OLAP engines.
How explicit producer–consumer contracts cut pipeline incidents by 60% across hundreds of teams.
Bridging the online/offline gap with a single source of truth for features and labels.
From dashboards to applications: how the cloud warehouse became a programmable platform.
A blueprint for feature stores that stay fast, consistent, and debuggable as model count explodes.
What it really takes to run replicated, strongly-consistent storage across continents.
Moving teams beyond correlation dashboards to causal models that actually inform decisions.
How cost-based optimizers are being rebuilt for Iceberg, Delta, and the disaggregated stack.
A practical framework for building evaluation suites that catch real regressions, not vibes.
Designing internal developer platforms that abstract Kubernetes without hiding what matters.
Serving large embedding models and fresh features with single-digit-millisecond latency.
Lineage, freshness, and anomaly detection — the telemetry that keeps data products reliable.
Where privacy-preserving training genuinely pays off — and where a central warehouse still wins.
Treating freshness, completeness, and contracts as SLAs your consumers can depend on.
A look back at how warehouses and lakes converged, and why open formats won the decade.
Lessons from operating shared ML infrastructure for distributed teams overnight.
Safe, backward-compatible schema change strategies for petabyte-scale tables.
Autoscaling analytical workloads without paying for idle capacity between the bursts.
The opening keynote that launched the conference: separating storage from compute, for good.
Replacing brittle scheduled ETL with change-data-capture and streaming so account updates propagate to every downstream system in seconds, not hours.
How governed, self-serve platforms let analysts ship fast while keeping lineage and access under control.
Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.
Designing A/B systems that account for interference, peeking, and the messy reality of product metrics.
Keeping vector indexes consistent and low-latency as embeddings and documents change constantly.
Turning raw billing exports into engineering decisions developers can actually act on.
Making the common ML workflows boring: paved roads that turn heroics into routine releases.
Keeping operational databases and the lakehouse in sync without melting the source system.
The tooling, runbooks, and incentives that keep data products dependable and engineers rested.
Why most ML benchmarks mislead, and how to build evaluation suites that survive contact with reality.
Design tradeoffs in LSM trees, replication, and consensus for databases that span continents.
Connecting data platforms across providers without opening holes your security team will hate.
Translating fairness research into measurable guardrails teams can actually put into production.
Batch, streaming, and CDC under one roof — choosing the right tool for each source without sprawl.
Geo experiments, switchbacks, and synthetic controls when a clean A/B test just isn't possible.
Org design, interfaces, and the platform-as-product mindset behind ML infrastructure that scales.
Defining freshness and completeness SLAs your consumers can plan around — and meeting them.
Backpressure, schema drift, and dead-letter handling for ingestion pipelines that stay calm.
Hard-won lessons building approximate nearest-neighbor systems for recommendations at scale.
Right-sizing compute, tuning file layouts, and killing the silent spend in your data platform.
Reproducible, shareable evaluation pipelines for teams that no longer sit in the same room.
A practical tour of Raft, Paxos variants, and the failure modes that bite real deployments.
When spanning providers genuinely pays off — and the networking realities nobody warns you about.
Lessons from standing up governed, remote-first data platforms when the office disappeared.
Batching and model selection strategies from the days when every GPU hour really hurt.
What recurring incidents taught one team about lineage, alerting, and blameless postmortems.
Decoupling storage growth from compute cost when workloads are spiky and unpredictable.
Building the first in-house A/B platform: metrics, guardrails, and the trust it took to earn.
An early call for rigorous, reproducible model evaluation as data moved into the cloud warehouse.
The storage-engine decisions behind separating compute from durable, columnar data.
How model-assisted automation reduced manual reconciliation and incident toil across a multi-account, event-driven integration estate.
Policy-as-code, automated guardrails, and infrastructure templates that keep enterprise cloud estates compliant without slowing delivery.