Data

Your CDP Just Got Its Own Brain. It Never Leaves the Building.

Marie Vackova
Marie Vackova Marketing Operator
Your CDP Just Got Its Own Brain. It Never Leaves the Building.

Every marketing team running AI on customer data has the same conversation eventually, and it usually happens in a meeting nobody scheduled, when somebody from security asks where the data actually goes.

The honest answer is uncomfortable for most teams. It goes to an API, to a provider whose terms say they will not train on your raw data with an asterisk attached, somewhere you cannot audit, under a bill you cannot predict.

Meiro can now deploy a tested, auditable open-source model directly alongside your customer data, in your cloud, your region, your provider, or entirely on-premise.

Why this became possible now

Open-source models have always trailed the frontier labs, and the received wisdom puts the gap at three to six months depending on which benchmark you trust. That gap used to matter enormously and it matters much less now.

Think about what a frontier model could do eight months ago. It was remarkable, and that same capability is available as open source today, which means it can be deployed on infrastructure you own, audited by people you employ, and closed off from the outside world.

That shift is what let us pack the whole thing into one deployment: data collection, identity resolution, the journey canvas, segmentation, communication, and now the intelligence that drives them.

The part most teams get wrong about PII

Most conversations about safe AI stop at redaction. Strip the email address, mask the phone number, ship the rest, problem solved.

It is not solved. For an AI assistant to be genuinely useful it has to run queries and analyse the results, and even aggregated, anonymised data tells a capable model an enormous amount. The intelligence is powerful enough to fill in the gaps, which makes redaction a speed bump rather than a wall.

So we approached it from the other end. Meiro embeds a small, specialised model that runs real-time PII detection inside the application itself, fast enough and cheap enough to run on the setup you already have.

It catches what pattern matching misses. An address, a first name, a surname and a date of birth sitting together become identifiable even though no single field is, and the same goes for cookies, user IDs and marketing identifiers. The model recognises the shape of identity rather than its obvious markers, which lets us draw a boundary at every exit point: what is displayed to a user, what is exported to a destination, what goes into a custom audience or out to a call centre.

Capable enough beats most capable

There is a subtext here worth saying out loud, which is that you rarely need the most capable model available. You need one capable enough for the task in front of you, without the trade-off of not controlling it.

Journey decisioning is the clearest example. Choosing the next best action for a customer does not require a flagship frontier model, and a small specialised one does the job for almost nothing, fast enough to operate at scale, often with better results for that particular decision than a general-purpose giant would produce.

The assistant layer is where the recent progress really shows. Piper, our assistant across Pipes and Engage, can now follow genuinely long-horizon tasks: building campaigns, generating banners, analysing your data and telling you where the gaps are. Because we supply a curated set of skills on top of the intelligence, Piper can reason about what you are not doing yet.

I can see you are running abandoned basket and shopping intention. I can also see you have a connected product feed and the data to run price drop. Here is the email template, here is the journey canvas. Should I run it?

Piper proposes and a person approves, and everything that happens inside the platform is audited.

How we know good enough is actually good enough

Claiming a smaller model is sufficient is easy, and proving it is the work.

Evals are to AI what tests are to software. You change something and you verify that nothing broke. We have built a broad evaluation suite out of the real conversations our analysts have while configuring the platform, distilled into complex, representative tasks.

Each model gets dropped into the harness where the AI actually lives, with its tools, context, skills and data, able to write queries, execute them and read the results. We swap in one model at a time, change nothing else, and watch what happens across hundreds of use cases.

When a new open-source model appears, and one appears almost weekly now, we do not wait for a frontier lab’s roadmap. We run the suite and find out precisely where that model is more than good enough.

What it costs to run

Our current model of choice is Qwen3.8-27B, an open-source model with 27 billion parameters, running on 32GB of VRAM. Frontier models are now measured in trillions of parameters, so this is a comparatively small one, and it performs very well in our benchmarks.

In practice that is a graphics card, or a box costing a few hundred dollars a month. It is the current sweet spot between cost of operation, model size and what the thing can actually deliver, and if you need more capability it scales, with the eval suite telling us what you would gain.

Which brings up the economics, because you are not paying per token, you have a fixed cost. There is a real debate about where AI pricing goes from here, cheaper as compute improves or more expensive as today’s subsidies unwind, and you do not want to discover the answer through a bill that is a hundred times what you budgeted for what has quietly become the most important software you pay for.

Three reasons to take this seriously

Your data never leaves your environment. Contacts, prompts, queries and results all stay inside. You are not pushing any of it to a third party whose terms say they will train on it but not the raw data, and you are not surrendering the asset.

You get control and auditability. Open models can be inspected, the deployment is yours, and every action inside the platform is logged. Some of our clients already run their own hosted models with a thin layer on top for cost accounting and visibility into what their users are sending.

Your costs become predictable. Fixed infrastructure rather than variable per-token exposure.

Where this goes next

Bundling the model with the CDP is one deployment mode among several. The platform is agnostic about where intelligence comes from, whether that is Gemini, OpenAI, Anthropic, a local provider or your own API endpoint, and if you already run a model in-house you can tell us the API and we will connect to it.

Once that in-house intelligence exists it stops being a CDP capability and becomes a company-wide one.

Model choice will keep changing, which is why the system is built to swap models and why we treat keeping pace as our job rather than yours. You should not be spending your headcount tracking model releases.

Customer experience is one of the few frontiers brands will genuinely control from here, and the data underneath it is not something to give away.

See how a local AI CDP works for where the model runs, and AI agents in Meiro for what Piper does with it.

See where the model runs

On-premise or inside a private cloud account, in the same boundary as the data it reads.

Marie Vackova

Marie Vackova

Marketing Operator

Marie is a Marketing Operator at Meiro, focused on content for the CDP and data infrastructure side of the product. Background in enterprise tech and consumer brands, across the GCC and Europe.