What CustomerLake Is: An Agentic CDP Built Into the Lakehouse
Databricks CustomerLake is an AI-native, agentic CDP platform that embeds identity, segmentation, and activation workflows directly into a customer data lakehouse, keeping governed data, AI models, and marketing agents together in one environment for continuous, real-time personalization at scale. Databricks positions CustomerLake as more than a new martech tool: it is a way to move marketing execution to where enterprise data and machine learning already live. Instead of exporting data into a standalone CDP, CustomerLake operates inside the Databricks Lakehouse, using Unity Catalog for governed access and integrating into the wider advertising and marketing stack. Ali Ghodsi, Databricks’ co-founder and CEO, describes the shift as moving from campaign calendars to “agents that constantly analyze, decide, and act on every customer in real time,” enabling what the company calls “infinity campaigns” and 1:1 personalization “a billion times a day.”

Eliminating CDP Data Silos With a Customer Data Lakehouse
CustomerLake directly targets the long-standing data fragmentation that limits many enterprise CDP architectures. Traditional stacks spread customer profiles, events, and consent records across analytics warehouses, marketing clouds, and channel tools, with each system holding a slightly different version of the truth. That fragmentation weakens AI models, complicates governance, and forces teams to spend time copying and reconciling data instead of improving customer journeys. By placing CDP functions on top of the existing customer data lakehouse, Databricks removes the need to pipe copies into a separate environment. Identity resolution, profile management, segmentation, and activation all run against a single governed data foundation, shared with finance, product, and operations. This reduces latency between insight and action and limits the storage bloat and reporting mismatches that come from duplicative datasets, while also giving compliance teams a clearer view of where customer information resides.

From Campaign Calendars to Agentic, Always-On Personalization
Databricks frames CustomerLake as a break from campaign-centric workflows built around batch cycles: plan, build segments, launch, then measure. In this AI-native CDP model, autonomous and semi-autonomous agents form a real-time personalization engine that constantly monitors behavior, determines offers, selects channels, and times messages for each individual. CustomerLake introduces “campaign agents” and “profile agents” that can draft briefs, assemble audiences, resolve identities, and orchestrate activation as a continuous loop rather than a sequence of tickets. Teams can keep “humans in the loop” by approving recommendations before agents act, and then gradually expand automation as trust grows. By collapsing analysis, decisioning, and execution into one agent-driven system, the platform aims to support 1:1 experiences at massive scale while still respecting enterprise needs around approvals, audits, and policy controls.
Governed AI, Open Integrations, and Enterprise CDP Architecture
CustomerLake’s architecture ties governed data and AI models together instead of pushing decisions into separate execution tools. Unity Catalog controls which teams and agents can access which customer attributes, while the lakehouse stores both raw data and refined profiles for AI-native CDP use cases. Databricks describes an open ecosystem: CustomerLake can call third-party models or agentic systems via APIs or Model Context Protocol, and it connects to identity partners such as Acxiom, Epsilon, LiveRamp, TransUnion, and Adstra for enrichment. On activation, it supports major ad and marketing platforms including Adobe, Meta (with Conversions API), The Trade Desk, Braze, Iterable, Snapchat, and more. This keeps the real-time personalization engine close to governed data while still using existing martech and adtech investments, reshaping enterprise CDP architecture around the lakehouse rather than standalone hubs.
Private Preview and the Road to Enterprise Adoption
CustomerLake is currently in Private Preview, signaling that Databricks is targeting complex enterprise CDP requirements before a broader release. Early adopters named include HP, Circle K, AB InBev, and Getnet by Santander, all of which already rely on the Databricks platform for analytics and AI. For these organizations, an agentic CDP platform inside the lakehouse promises less integration overhead, fewer data copies, and more consistent governance compared with maintaining separate CDPs and activation tools. The focus on Private Preview also reflects the need to validate “always-on” agent workflows, human approval patterns, and real-time personalization use cases in demanding production environments. If those pilots prove successful, CustomerLake could accelerate a wider shift from batch-driven CDP deployments toward lakehouse-native, agent-driven decision systems that blur the line between data infrastructure and marketing execution.






