All conferences

Edition 9 · 2025

Intelligence on Every Table

San Francisco, CAMoscone Center

July 21–23, 2025

11,500

Attendees

168

Talks

63

Papers

Headliners

Keynote speakers

Program highlights

Keynote

Reading the Cloud Bill: FinOps for Engineers

Turning raw billing exports into engineering decisions developers can actually act on.

Ben Fisher Jul 21, 2025 · 2:30 PM
Keynote

The Lakehouse Decade: Where Storage Meets Intelligence

A decade-long look at how open table formats reshaped the modern data stack and what comes after the lakehouse.

Dr. Amara Okafor Jul 22, 2025 · 12:30 PM
Keynote

Real-Time Sync: Event-Driven Integration Across Enterprise Systems

Replacing brittle scheduled ETL with change-data-capture and streaming so account updates propagate to every downstream system in seconds, not hours.

Srikanth Jonnakuti Jul 23, 2025 · 9:00 AM
AI / ML

Golden Paths for Production ML

Making the common ML workflows boring: paved roads that turn heroics into routine releases.

Anita Desai Jul 21, 2025 · 9:00 AM
Data Engineering

Consensus at Planet Scale

What it really takes to run replicated, strongly-consistent storage across continents.

Kenji Tanaka Jul 21, 2025 · 11:00 AM
Analytics

Experimentation at Scale Without Fooling Yourself

Designing A/B systems that account for interference, peeking, and the messy reality of product metrics.

Hana Kim Jul 21, 2025 · 1:00 PM
AI / ML

Embeddings in Production: Retrieval That Stays Fresh

Keeping vector indexes consistent and low-latency as embeddings and documents change constantly.

Lucia Fernandez Jul 21, 2025 · 2:30 PM
Cloud

Secure DevSecOps for Multi-Account Cloud Infrastructure

Policy-as-code, automated guardrails, and infrastructure templates that keep enterprise cloud estates compliant without slowing delivery.

Diego Ramirez Jul 21, 2025 · 4:30 PM
AI / ML

Building Feature Platforms That Scale to Thousands of Models

A blueprint for feature stores that stay fast, consistent, and debuggable as model count explodes.

Priya Nair Jul 22, 2025 · 9:00 AM
Data Engineering

Self-Serve Data Platforms That Don't Break Trust

How governed, self-serve platforms let analysts ship fast while keeping lineage and access under control.

Naomi Clarke Jul 22, 2025 · 11:00 AM
AI / ML

Tail Latency Is the Product: Serving LLMs in Real Time

Batching, speculative decoding, and scheduling tricks that keep p99 latency low under bursty load.

Arjun Patel Jul 22, 2025 · 11:00 AM
Cloud

Intelligent Automation for Cloud-Native Integration Platforms

How model-assisted automation reduced manual reconciliation and incident toil across a multi-account, event-driven integration estate.

Srikanth Jonnakuti Jul 22, 2025 · 3:00 PM
AI / ML

Serving Trillion-Token Models on a Budget

Practical patterns for batching, quantization, and autoscaling LLM inference without burning your cloud bill.

Rohan Mehta Jul 23, 2025 · 3:00 PM

Proceedings

Peer-reviewed papers — 2025

Delta Lake Time Travel for Point-in-Time Analytics at Petabyte Scale

A. Okafor, J. Rivera

Data Engineering202523 citations

Self-Improving Training Pipelines with Active Data Selection

R. Mehta, V. Cruz

AI Systems202518 citations

Adaptive Partitioning for Petabyte-Scale Lakehouses

A. Okafor, M. Bell, J. Rivera

Storage202523 citations

Cost-Aware Scheduling for LLM Inference Clusters

R. Mehta, V. Cruz

AI Systems202519 citations

Contextual Embedding Caching for Low-Latency RAG Inference

A. Patel, M. Bell

AI Systems202512 citations

Schema-Aware Data Mesh Catalogs with Automated Lineage Discovery

N. Clarke, D. Chen

Data Engineering202516 citations

Windowing Strategies for Late-Arriving Data in High-Throughput Streams

L. Zhao, K. Nakamura

Streaming202521 citations

Retrieval-Augmented Generation Pipelines Over Proprietary Document Corpora

R. Mehta, X. Dubois

AI Systems20258 citations

OpenTelemetry-Based Observability for Serverless Data Functions

O. Müller, W. Kim

DevOps20256 citations

Tiered Storage Compaction Strategies for Time-Series Databases

U. Johansson, M. Bell

Data Engineering202513 citations

Decentralized Data Mesh Governance with Federated Computational Policies

N. Gupta, Y. Patel

Data Governance20254 citations

WASM-Based User-Defined Functions for Portable Stream Processing

I. Santos, Z. Thompson

Streaming202510 citations

k-Anonymity Guarantees in Real-Time Data Publication Pipelines

P. Adeyemi, D. Chen

Security20259 citations

Multi-Modal Embedding Pipelines for Unified Search Across Data Assets

V. Cruz, S. Volkov

AI Systems202515 citations

Continuous Integration for Data: Testing Schema Changes in CI/CD

Q. Zhang, O. Müller

Data Engineering202512 citations

Just-in-Time Data Transformation with Query-Time Materialization

W. Kim, C. Williams

Data Science20257 citations

Stateful Serverless Functions with Durable Exactly-Once Guarantees

J. Rivera, B. Andersen

Cloud202511 citations

Graph-Based Dependency Resolution for Polyglot Data Pipelines

T. Han, F. Almeida

Data Engineering20256 citations

Predictive Cost Modeling for Multi-Cloud Data Platforms

B. Fisher, E. Larsen

Cloud202517 citations

Liquid Clustering for Adaptive Layout Optimization in Data Lakes

D. Chen, R. Mehta

Data Engineering20254 citations