// Field notes from the founders

Why the next wave of AI is about structured data — a D2C take

Julian Aylward · Co-Founder / Chief Data OfficerJune 25, 2026Originally published on Medium
Ask your Claude why retention dipped last month and you'll get a beautiful answer. Structured. Confident. Authoritative. A paragraph (or 10) you could paste straight into a board update.The question is, well what did you put into Claude in the first place? And how trustworthy is it? To which the answer is normally 🤷This isn't an isolated problem. This happens day in, day out at most D2C brands, who run on metrics that are quietly wrong with the real problems and opportunities drowned out by the noise. We've seen this inside brands far more sophisticated than people assume.To be clear, this is not a new problem, it's been here as long as I can remember. Over a decade in the trenches as an analyst means highly sensitive alarms immediately start ringing and I instinctively reach for Twyman's Law (Any figure that looks interesting or different is usually wrong).CAC is lifted from inside Meta, Google and TikTok with each platform proudly claiming credit for the same conversion. Different attribution methods and lookback windows create a mirage of numbers. There is no clean split between a genuinely new customer and a repeat/reactivated one based on a cookie. Add the three numbers up and you've triple-counted your way to a cost-per-acquisition you can't actually act on. Then there is the metric itself. Whenever someone says "ROAS", it makes me want to scream. I could write an entire blog post (or book) about my hatred of ROAS.LTV is a flat assumption that hardened into fact years ago: "the average customer places about four orders." Nobody's rebuilt it from a real cohort curve since. It sits in a static spreadsheet updated monthly for the board meeting with a magic multiplier in cell G4 that gets you from 3 months LTV to 12 month LTV. It is retrospective and not predictive, and distorted by seasonality and early stage discounts. Is December's cohort really so bad? By the time you find out, the ship will have sailed ⛵Retention — the single most important number in a subscription business is a spreadsheet the Head of Growth refreshes when she remembers to. Monthly if you're lucky. How is it defined? We'll never know. We've seen mature, well-run brands operate on this fiction for years before we finally fixed it.Until now at least, in many cases AI is again making this worse, not better. The tools encourage uploading or attaching spreadsheets, connectors are springing up that make integrations easy with tools like Google Ads, Google Analytics etc but uptake is mixed. Few companies and fewer individuals are connected up to their backend, or data warehouse and have the foundations in place to make this trustworthy and safe.Companies who don't have the foundations in place have twitchy IT managers and data teams. It's not safe, the data isn't sufficiently well structured, the metrics are not defined, "we need to build a proper semantic layer", and in many cases their concerns are not without justification.This leads to people "plugging in whatever they can", rather than "the things they really need".The data team (if you have one) can act as "guides" through the quagmire, but it typically requires a seasoned data team to build the data models, semantic layers and agentic interface that enables this to scale. Without this, models don't fix bad inputs, they launder them. It takes your four-orders assumption and builds a cogent, confident, paragraph-long recommendation on top. Confident wrong is far more dangerous than obviously wrong, because you act on it. You reallocate £40k from Google to Meta off the back of it, you miss the fact your retention cohorts are sagging and LTV is cratering due to ramping up acquisition of low retention, low margin customers.

There is a quiet consensus forming around this

There's a real shift happening in how serious people think about AI, albeit somewhat under the radar and it's worth a D2C operator knowing about. The frontier has moved from "who has the cleverest prompt" to "who feeds the model the right, structured context." The data world calls it the semantic layer; the broader term is context engineering. Gartner now tells data teams to pivot toward exactly this.This is a view also recently aired by the Databricks CEO. You might think that he has a vested interest here, and he does, but once you say it out loud, it seems obvious.And yet... as of June 2026 (let's see how fast this ages), most D2C brands are not able to feed in robust metrics from trustworthy sources that allow them to take proper advantage of AI's rapidly increasing capability to leverage (high quality) structured data. No matter how good the current class of LLMs become, they cannot solve for data or context they do not have.This is something that we data folk have believed from day one. Asking the right question is a challenge in and of itself, but even in the hands of a seasoned operator, the opportunity is squandered if your inputs are wrong. S**t in, s**t out.

What "right inputs" actually means for a consumer brand

If the 6 month, £500k [re]build a data team and data warehouse isn't already on your roadmap for the year, then what do you do? What about solving for just the eight-to-ten numbers that genuinely move profit, correctly, in one shared place: true blended CAC split by new versus reactivated; LTV built from real cohort retention curves or a predictive model, AOV and contribution margin; payback per cohort.There are places where the raw data is genuinely messy, or just absent. Pick/pack/ship costs are the classic, they vary by SKU, channel and 3PL and are often invoiced months late and riddled with rebates and adjustments. In this case, you don't "refuse to go there", you standardise the estimate, expose and log the assumption, and make "good enough" work safely for you. A defensible estimate today beats an over-engineered, brittle solution that will start drifting from reality as soon as you build it, or doing nothing at all.Get the inputs right and AI finally earns the confidence it's been faking. Same model, same prompts, but now the brilliant answer is built on a true number instead of whatever it had to hand at the time.That's the unglamorous half of the "company brain" nobody puts on a slide, but we know that it needs fixing before real value can be unlocked.Without this, you cannot pass go and you will not collect £200.Check out Signal over Noise to apply for the Pilot!