Customer MDM Intake: Where Data Quality Actually Dies
Jan 2026 · Data GovernanceEvery bad customer record gets blamed on the same thing: "the MDM platform isn't matching correctly." I've sat in enough of those conversations to say it plainly — the platform is usually innocent. By the time a record reaches the master data hub, the damage is already done. I led a rapid process analysis of customer master data intake at a global pharmaceutical company: how customer data got acquired, validated, and stored before it ever reached the SAP MDM and Oracle environment sitting downstream. What we found wasn't a matching-algorithm problem. It was an intake problem, and it had been for years.
Everyone Blames the Platform
The pattern is consistent across organizations: duplicate customer records pile up, golden-record merges produce nonsense, sales and service teams stop trusting the "single view of the customer" the MDM program promised, and the conversation turns to tuning match rules or swapping platforms. That conversation is usually a year too late and pointed at the wrong layer. An MDM platform matches, merges, and survives records using rules applied to whatever data actually arrives at its door. If what arrives is inconsistent, the platform's job is to make the best of bad inputs — and "the best of bad inputs" still produces bad golden records.
The engagement started as a data-quality investigation and became, within the first round of stakeholder interviews, an intake investigation instead. Nobody had mapped, end to end, what happened to a customer record between "someone typed something into a form" and "record lands in the master data hub." Once we mapped it — acquisition, validation, storage, each as its own stage with its own owners and its own failure modes — the anomalies in the golden record stopped looking mysterious. They looked inevitable.
Why Process Analysis Finds What Profiling Tools Can't
Data profiling tools are good at telling you what's wrong: null rates, format inconsistencies, duplicate clusters, orphaned foreign keys. They're bad at telling you why, and why is the only thing you can act on. Profiling a table shows you that 30% of a phone number field is malformed. It doesn't show you that three different regional teams enter phone numbers in three different formats because nobody ever built a single intake form, or that one team pastes numbers from a CRM export that's never been reconciled against the live system.
That's what stakeholder interviews and process mapping are for. We ran structured interviews across sales operations, customer service, and the data stewardship function, then built Visio-level process maps of the actual acquisition-to-storage flow — not the flow in the training documentation, the flow people actually followed, including the workarounds. The gap between those two things was the finding. Documented process said one thing; the day-to-day reality, shaped by system limitations and deadline pressure, was another. Process analysis surfaces that gap. A profiling tool, looking only at the data that landed in the table, has no way to see it.
The Intake Anti-Patterns
A handful of patterns showed up repeatedly, and I'd bet money they show up in most customer MDM programs built up over a decade of acquisitions and system patches:
- Free-text fields feeding match rules. Company name, address, and contact fields captured as unconstrained free text, then fed directly into fuzzy-match logic downstream. Every typo, abbreviation, and "Inc." vs. "Incorporated" vs. blank becomes a matching miss the platform gets blamed for.
- Validation that lives in people's heads. The most experienced reps in each region knew, informally, which fields "really" needed to be right and which ones the system would accept garbage in. That knowledge never made it into a system rule — it lived in tenure, and it left when they did.
- Duplicate-creation paths nobody owns. Multiple entry channels — direct sales entry, self-service portal, batch import from an acquired business unit — could all create a "new" customer with no real-time check against what already existed. Each channel had an owner. The overlap between channels didn't.
- "We'll fix it in the golden record." The most damaging pattern, because it sounds reasonable. Teams knowingly let bad data through at intake because "MDM will merge/survive/clean it up." Survivorship rules can pick the best of several bad records; they cannot manufacture a good one that was never captured.
Designing the Front Door, Not a Better Mop
The recommendation set out of the engagement wasn't a platform change. It was upstream: consolidate the entry channels into a streamlined intake portal, move validation logic out of individual employees' heads and into the form itself, and give every entry channel a real-time duplicate check instead of an after-the-fact merge. We built a Figma wireframe of what that portal could look like — required fields enforced at entry instead of discovered at audit, a duplicate-check step before a new record could be created at all, structured fields replacing the free-text ones feeding match logic.
The wireframe mattered more than it sounds like it should, because it made the recommendation concrete enough for stakeholders to react to. "Improve data validation" is a slide nobody argues with and nobody acts on. A wireframe showing the exact moment a duplicate check would have caught a specific record type generated actual pushback and actual buy-in in the same meeting — people could see precisely which workaround the new design would close, and which of their team's habits it would break.
That's the reframe worth holding onto: fixing customer master data by improving MDM survivorship rules is designing a better mop. Fixing it by redesigning intake is fixing the leak. One of them is a lot less glamorous to present to leadership and a lot more durable. It also plugs directly into how a stewardship function has to operate day to day once the new intake process exists — the workload of enforcing those rules doesn't disappear, it moves.
SAP MDM, Oracle, and the Limits of the Platform Layer
The environment here ran SAP MDM with source data touching an Oracle environment upstream — a common combination in large enterprises that grew through acquisition, where the "system of record" is really several systems of record stitched together over time. None of the intake findings were specific to that stack; they'd show up under any MDM platform, because they're upstream of the platform boundary by definition. What the specific stack did determine was where the validation logic could realistically live — some checks belonged in the intake portal itself, some had to live as staging-layer rules before data hit SAP MDM, and getting that split wrong (validating too late, or duplicating validation in three places) was its own source of maintenance drag.
If your organization is also thinking about how a data catalog or dictionary should describe these customer fields so the intake fixes stick over time, that's a related but distinct problem. And if any part of your intake process is starting to involve AI agents pulling or enriching customer data before it lands in MDM, the access-control questions compound fast.
Nobody budgets for a data quality problem that's still one system upstream of where the dashboards say it lives. Find the intake process before you fund the platform fix — the platform was never where the damage happened.
What Would Make This Wrong
- If your bad data genuinely originates from an external, uncontrollable source — a third-party data broker feed, for instance — redesigning your own intake portal won't fix it. The fix in that case is a validation layer at ingestion, not a front-door redesign.
- Portal redesigns take real adoption time. If the organizational appetite for a new mandatory workflow is low, the free-text-and-fix-later pattern can persist even after a better form exists, because the old habit is faster in the moment.
- Process mapping via interviews captures what people say they do and what a sample of cases showed us — not a full statistical census of every entry path. Large organizations should validate map findings against a broader data sample before committing capital to a redesign.
- Some validation genuinely belongs at the platform layer (cross-system duplicate resolution at scale, for instance) — pushing everything upstream into intake isn't free either, and the right split is workload- and architecture-specific.
Related reading: Data Stewardship Rollout: What Month 3 Actually Looks Like · Data Dictionaries and RAG-Based Governance · Governing AI Agents' Data Access
Data & AI governance, from the field
Notes on governance that actually ships — and the free RFP & TCO scorecard as a welcome gift.
No spam, unsubscribe anytime.
You're in — check your inbox for a welcome note.