The Agentic CDP is Coming to Your Data Platform. Is Your Customer Data Ready?
TL;DR
- Databricks has launched CustomerLake, a CDP where AI agents build customer profiles and run campaigns for you. It sits inside the lakehouse, next to data you already hold.
- It is in private preview, with general availability expected in late 2026 or early 2027.
- CustomerLake works from data that is already in your lakehouse. It does not collect web or app behavior itself.
- So the quality of your event data sets a limit on how well the agents perform. Thin or messy behavioral data produces thin profiles, however good the agents are.
- Before you migrate, test your collection layer against the four things Databricks says an AI stack needs: context, control, cost, and choice.
"AI does not have an intelligence problem. It has a context problem."
Ali Ghodsi, CEO & Co-founder, Databricks
These were the words from Ali Ghodsi at this year's Databricks Data + AI Summit. And we couldn’t agree more. You can put the best model in the world behind your customer data, but if it can't see what a customer is doing on your site right now, and everything they did before this visit, it has nothing useful to say about them.
At the same event, Databricks announced CustomerLake, an agentic CDP built inside the lakehouse. This new generation of CDP provides you with agents that build customer profiles and run engagement continuously, sitting next to the data you already have, under the governance you already wrote.
This makes perfect sense and we firmly believe this is the direction CDPs are heading.
CustomerLake is a separate product, bought separately, and it will not be generally available until late 2026 at the earliest. But plenty of Databricks customers are already expressing interest.
This comes at a time when Gartner is already advising CMOs to reassess their data strategy before renewing a CDP contract into 2027. And whether your answer turns out to be CustomerLake, a rival agentic CDP, or the stack you already run, one question decides how much you get out of it.
How much of your customer context can the agents actually see?
Agents can only act on the context they’re given. Ask one of them to flag customers at risk of churning and it will do a pretty good job with what it can see. But if your event stream only carries completed purchases and page views, the agent can’t see the search that returned nothing, the pricing page someone bounced off, or the checkout that threw an error on the third attempt.
That holds for both kinds of agent you’re running. The ones that answer questions, like Genie, and the ones that take actions, like Campaign Agents. An analyst asking why churn is up and an agent deciding what to send a customer are reading the same context, and they’re limited by it in the same way.
There is a bigger version of that point. The context an agent needs is the same context your lakehouse needs, your models need, and your real-time systems need. Snowplow has been building that single layer since inception, feeding data platforms, lakes, streams, and now agents. Agents have just made this requirement impossible to ignore.
So before diving into the agentic CDP world, you need to fully assess your customer behavioral data.
In this piece, we’ll cover what CustomerLake is, why the category is moving into the data platform, and the four tests to run on your collection layer before you migrate.
What is an agentic CDP?
An agentic CDP is a customer data platform where AI agents act as the primary operators. They build profiles, assemble audiences, choose the next action, and adjust as results come in.
The reasoning behind the shift is worth following, because it explains why the architecture changed and not just the interface.
Buyers increasingly consult an LLM before purchasing, and will increasingly delegate the purchase itself, so decisions now resolve in seconds rather than weeks. If we think about a traditional campaign, it usually follows a waterfall approach. So you have your strategy, then a customer 360 assembled by data engineering, then audiences built by marketing ops, then creative, then activation. All of which takes weeks, which is fine against a buyer who also takes weeks.
The alternative to this approach is continuous personalization at the individual level. This means you have an agent per customer rather than a campaign per segment. An agent deciding in the moment needs the data and the context where they already live, which many traditional CDPs simply cannot provide.
That is why agentic CDPs are moving towards the data platform. This also raises the stakes on what the agent can see, because a loop that closes in seconds has no overnight batch job to fall back on.
Some agentic CDPs like CustomerLake are embedded in the data platform. Others are managed platforms holding their own copy of your data.
The direction is settled
The debate around where an agentic CDP should run is pretty much settled. Gartner expects that by 2030, 80% of net-new enterprise CDP deployments will be embedded in, or composable with, data platforms rather than standalone. What’s driving this is duplication. A separate CDP means a second copy of customer data to govern, pay for and keep in sync, and an agent deciding inside a session cannot wait for that sync.
Gartner is describing where this ends up. What standalone CDPs are doing in the meantime says the same thing from the other direction.
The day after CustomerLake was announced, BlueConic acquired Blueshift, with the stated aim of giving AI agents real-time behavioral context to decide and act on. That is a CDP buying agentic capability rather than sitting inside a platform that already has it. And two of the best-known independents have already stopped being independent. Uniphore bought ActionIQ in December 2024, and Rokt bought mParticle six weeks later.
None of the three is moving into a data platform. They are consolidating and buying to keep pace, which is what the standalone model looks like under strain.
David Raab, who founded the CDP Institute, read the launch as “a good move by Databricks, moving up the value chain from data management to execution,” and “a challenge for many conventional CDP vendors.” Gartner also expects Snowflake and others to follow Databricks into the agentic CDP space.
Tasso Argyros, who founded ActionIQ, put it plainly in an Adweek interview alongside Ghodsi: "I think the CDP, as middleware, is going to go away."
Interestingly, the CustomerLake team is led by Argyros and former product lead Justin DeBrabant, working from a New York team with a lot of ex-ActionIQers in it. This is the very team that built the enterprise CDP category and it is now rebuilding it inside a data platform. We think this shows how seriously the bet is being taken.
We hold the same view about where this ends up. The data platform really is the center of gravity for storing customer data, and a CDP holding its own copy alongside it was always going to be a temporary arrangement.
What Databricks CustomerLake actually is
Let’s look a little closer at CustomerLake itself. Databricks markets it as "the Agentic CDP built in Databricks." It is now in private preview with the likes of HP, Circle K, AB InBev, and Getnet by Santander.
Within the solution, two agent families do the work. You have Profile Agents that turn raw records into a governed Customer 360. These agents resolve fragmented identities through Agentic Identity Resolution, where agents generate the matching rules, LLMs take over when those rules are inconclusive, and it falls back to human review when needed. The agents learn from resolved identities as they go. And then you have Campaign Agents, which run what Databricks calls Infinity Campaigns, continuous loops that react to customer context in real time rather than firing on a calendar.
Marketers can query it through Genie in natural language, so inventory levels or support history are in scope without anyone modeling them as traits first. Everything stays governed by Unity Catalog. And Activation goes out through partners such as Braze, Acxiom, and The Trade Desk.
The architecture is very much deliberate. CustomerLake starts where your data lands, and no client-side SDKs have been announced, because profiles are built from what is already inside the lakehouse.
So what lands in the lakehouse sets the ceiling.

Golden Context is the right idea, but something has to feed it
Golden Context is Databricks’ term for what a Golden Record leaves out. In simple terms, a Golden Record tells an agent who someone is, whereas Golden Context adds what’s going on.
Argyros, Ghodsi and Reynold Xin describe it as adding what the business is trying to accomplish right now, and what has already been tried with this customer and how they responded.
Look at that second ingredient. What was tried, and how they responded, is behavioral history. It is a record of things a customer did, in order, over time. It does not originate in the CDP. It arrives from wherever your events are captured, and CustomerLake reads it out of the lakehouse like everything else.
So Golden Context has a supply chain, and you own most of it.
Profile Agents infer traits from the records they are given. Agentic Identity Resolution reconciles the same customer across datasets you already own that do not share a customer key. What it cannot do is create an identifier that was never captured in the first place. If a customer's cookie expired after seven days, your lakehouse holds four separate anonymous strangers rather than one person with a four-week history, and there is nothing in those records for AIR to match on.
What a well-fed lakehouse looks like
Let’s take a look at Burberry, a joint Snowplow and Databricks customer. While Burberry may or may not be evaluating a CustomerLake today, we can take a look at how Burberry has been building customer context and delivering it to the lakehouse with a quick glance at their public numbers showcased in our joint case study.
With Snowplow and Databricks, the fashion retailer reduced clickstream latency by 99%. Cookie duration increased 52x, from seven days to twelve months, after moving to Snowplow's server-set first-party cookies. That data feeds 40 personalized models in Databricks covering product recommendations, propensity scoring, and lifetime value.
The 52x is the number to look at. Twelve months of continuous identity means twelve months of the customer journey, stored in the lakehouse and attributed to one person. Seven days means a two-week-old browsing session belongs to a stranger.
Point a Profile Agent at each of those and you get very different profiles out. The agent is identical in both cases, and so is the platform and the governance. The difference was decided upstream, before any of it reached Databricks.
Databricks named the four tests. Now apply them to your collection layer
Ali Ghodsi’s fascinating Summit keynote set out four key things enterprises need from their AI stack: context, control, cost, and choice. Ghodsi was explicit on control, saying teams need "security rules, policies, and auditability over the AI, the data, and the infrastructure." He was just as explicit on lock-in, calling out how organizations "keep adding to their complicated stacks."
Those four tests do not stop at the lakehouse boundary. If the data platform is your center of gravity, every tool feeding it has to pass the same four. Most companies have never asked these questions of their collection layer, but they really should.
Context: can an agent tell what happened?
Can you name the entities in your data? Products, subscriptions, baskets, plans. If your events are untyped JSON blobs, an agent has page views and clicks to reason over and not much else.
Is your data validated before it lands, or after? Validation at the point of collection is what lets an agent treat a profile as trustworthy. Cleaning up downstream in dbt works for a dashboard a human reads, and less well for a decision made inside a session.
Is identity resolved before the data lands, or after? Stitching anonymous activity to a known customer in-stream means every system downstream sees the same resolved person from the moment the event arrives. Leave it until after landing and the agent deciding right now is looking at an unidentified visitor, while every consumer that needs the answer works it out separately.
And is it available in-session? Agentic engagement happens while the customer is still there. Real-time decisioning needs a real-time collection layer behind it, or the loop only ever runs as often as your slowest hop.
Control: is the behavioral data in your lakehouse actually yours?
Most brands capture web and app activity somewhere. But GA4's native export only lands in BigQuery, and whatever arrives carries that vendor's schema rather than your own event design.
Unity Catalog governs what is in the lakehouse. It does not govern what a third-party SDK decided to collect, or how that vendor chose to structure it. If you cannot change the schema, you do not control the context.
Cost: what happens to collection when you move?
Licence fees are the easy number. Collection is the one that catches teams out, because an embedded CDP starts downstream of it. CustomerLake reads what has already landed in the lakehouse.
Which means the collection work happens either way. If your CDP handles tracking today, you re-instrument when you leave it. If you built ingestion in-house, it is usually welded to the stack you built it for. That bill arrives with the migration whether you planned for it or not.
The cost worth avoiding is paying it twice.
Choice: can you swap the agentic tooling without rebuilding the data?
Ghodsi's argument was about models. Locking an enterprise into one proprietary LLM creates platform dependency and architectural rigidity, and leaves you exposed to that vendor's pricing and deprecation decisions. Databricks' answer is to keep models swappable, routing complex reasoning to a frontier model and routine work to something faster and cheaper.
The same logic applies one layer down. CustomerLake will not be the only thing that wants your customer context. Your own product agents, your support tooling and whatever you build next all need the same behavioral history, and CustomerLake activates through Braze, Iterable, Twilio, Meta and The Trade Desk, so it was never designed to be the only consumer.
If that context only exists in a shape one platform can read, you have moved the lock-in rather than removed it. So the question is whether your collection layer can feed CustomerLake and everything else at the same time, without a second implementation.
How Snowplow fits with CustomerLake
We build the customer context, so what follows is our answer to our own four tests. Run the same questions at whatever else you are considering. You can also run this test yourself through a free 14-day Snowplow trial.
Context. You define your own events and entities in Event Studio, so a basket or a subscription arrives as declared structure rather than a blob an agent has to guess at, and Event Specification Validation checks each event against the spec you published before it lands.
Snowplow Identities then resolves identity in the stream. A stable snowplow_id is attached from a customer's first anonymous visit and stays with them through to login, so the record reaching your lakehouse is already stitched. That is the part AIR cannot do for you, because AIR matches identifiers that already exist. Hand it records that are already resolved and it has more to work with.
Signals then serves those attributes at sub-10ms p95, combining live in-session behavior with history from the lakehouse.
Control. The pipeline runs in your own cloud account. You own the schemas, the enrichments, and the destination. Unity Catalog governs your data once it lands. This governs how it was shaped before it got there.
Cost. You re-instrument once. While you evaluate CustomerLake, the same collection layer keeps feeding whatever you run today, so there is no parallel build and no cutover date to hit.
Choice. One collection layer, many consumers. The same validated events load Databricks, Snowflake, Redshift and BigQuery, forward to the marketing tools you already run, and serve your own applications and agents directly through the Signals API. CustomerLake becomes one consumer of your customer context rather than the place it lives, so adding or swapping agentic tooling later is a configuration change, not another re-instrumentation project.
What that makes possible
Customer 360. Your CRM and transaction tables are the Golden Record. Snowplow supplies the other half, what a customer is doing right now across every digital surface, so Profile Agents assemble both into one view rather than a purchase history with gaps. Supercell unified client-side web and server-side in-game behavior under shared identifiers, landing straight in Databricks.
Natural-language segmentation. Genie is only as good as the meaning attached to your columns. Snowplow enforces that meaning at collection through the schema registry, which is what turns "viewed this three times, never bought" into a question a marketer can ask directly rather than a ticket for the data team. At Strava, analysts and product managers self-serve from 4 billion events a day with no engineering support.
Consent you can prove. Consent reconstructed after the fact is an assertion. Consent captured as part of the event and carried into Unity Catalog policy is lineage, which is what a privacy team signs off on. Condé Nast lands every event validated, enriched and consent-compliant under one global consent setup, across 22+ brands and 12 markets.
Who this decision actually lands on
Gartner's advice to CMOs is the sharpest version of all of the above. Treat CustomerLake "as a data infrastructure decision rather than a traditional CDP procurement," and reassess your data strategy before signing extended contracts, especially any renewing in 2027 or later.
Raab made the related point that consolidating the CDP under corporate governance means "IT gains more control, which marketers may not appreciate."
If you read those statements together, it’s clear this is being written up as a marketing decision, but it is a data infrastructure decision, and the request to assess it will probably reach the data team before it reaches the CMO. If that’s you, the four tests above can act as a useful assessment.
Start building your customer context now
Collection is the one part of this you can fix today. It doesn’t require a wait on preview access, a GA date, or a decision about which agentic CDP you end up running. Whatever you choose, it reads from your lakehouse, and what is in there is already up to you.
If you want to see what validated, real-time behavioral data looks like landing in your own lakehouse, our free trial runs for 14 days in your own cloud account, with no sales call. Worth running alongside your CustomerLake conversations, because that context is what the agents will work from.
Disclosure: Databricks is an investor in Snowplow.