Companies start master data projects for two different reasons. The first is that the master data they already have is a mess — duplicates across ERP and CRM, conflicting values, nobody sure which record is real. The second is that the process for creating master data is a mess — a new supplier takes eleven days and three emails, and half the fields arrive wrong. These are not the same problem, and almost the entire MDM market is built around the first one.
I did not arrive at this framing on my own. It came from a practitioner in a data engineering forum making the point twice in the same thread, clearly having said it many times before: commercial MDM tools are really solutions for problem one, and they do not have much to offer for problem two. Once you have read that, you cannot unsee it in a vendor demo.
Two problems that look like one
Cleanup is retrospective. Records exist, they disagree, and you reconcile them. This is genuinely hard: working out that ABC Corp, A.B.C. Corporation and ABC Company Inc. are one entity is real engineering. Matching, survivorship rules, golden records — this is the machinery, and the tools that do it are good at it.
Creation is prospective. Someone in procurement needs a supplier that does not exist yet. What happens next decides whether you will be running another cleanup project in three years. Who can request it. What has to be captured. Who checks it is not already there under a different spelling. Who approves the payment terms. Which systems get it, and how.
At most organisations the honest answer to all five is some combination of a shared mailbox, a spreadsheet attachment, and one long-tenured person who knows the rules. That is the tap. Cleanup is the mop.
Why the market went one way
Not because vendors are lazy. Matching generalises. Fuzzy-matching a company name works roughly the same way at a bank and at a manufacturer, so it can be built once and sold repeatedly. It also demos beautifully — a duplicate count falling from 40,000 to 3,000 in a slide is a story a buyer can take to a budget meeting.
Creation workflow does not generalise, because it encodes your organisation. Who signs off a credit limit above fifty thousand is a question about your company, not about master data. Building for that means building something configurable and unglamorous, and it demos as a form with an approval box.
So the industry concentrated on entity resolution, and the authoring process got reimplemented from scratch inside every company that ever bought an MDM tool. Which leads to the observation that should worry anyone about to sign: an MDM hub that is only a destination is not used by very many people. The stewards do not live there. The requesters certainly do not. It becomes a system that receives data rather than one where data is made.
What happens when you only mop
Quality returns to roughly where it started, on a schedule set by how fast you create records. The deduplication project lands, the duplicate count drops, everyone is pleased. The intake process is unchanged, so it produces new duplicates at exactly the rate it always did. Two or three years later someone proposes a data quality initiative, and it is the same initiative.
The tell is a company that has done this twice. It is nearly always described as a data quality problem. It is a process problem wearing a data problem's clothes. We wrote about the mechanics of the decline in why your master data goes stale in four months.
What fixing the tap actually involves
None of this is exotic. It is mostly refusing to let record creation happen in unstructured places.
- One front door. Requests arrive in one place, not by email to whoever answers.
- Validation at entry. Required fields enforced while typing, not audited afterwards.
- Controlled values. Domain fields are chosen from a list. A country field that accepts free text will eventually contain nine spellings of the Netherlands.
- Duplicate check before creation. Search first. Catching it here costs one second; catching it in a match run costs a project.
- Approval where risk lives. Not every field — bank details and credit limits, not the description.
- Distribution from one point. Created once, sent to the systems that need it, instead of keyed into three of them.
Notice how little of that is matching. It is not a hard algorithmic problem. It is a workflow problem, which is why it rarely appears in a product built around entity resolution.
The question to ask in the demo
Vendors will show you matching, because matching is the impressive part. Ask instead: show me a person who is not a data engineer creating a new supplier, from request to it existing in the ERP.
If the answer routes through a professional services engagement, or a Power Automate flow you will build, or “that typically stays in your ERP”, you have a cleanup tool. That may be exactly what you need, if your problem really is a decade of accumulated duplicates. Just know which problem you are buying for, and do not expect the other one to improve on its own.
Which half Primentra is
Worth being direct, because this cuts both ways. Primentra is built for the creation half. Stewards work in a browser grid on your SQL Server, domain attributes are picked from controlled lists with cascading filters, changes can require approval before they go live, every field change is logged, and downstream systems read from integration views rather than being keyed separately.
What it does not have is a probabilistic matching engine. If you need to reconcile eight million customer records across a dozen acquired systems with fuzzy name matching and survivorship rules, that is Informatica or Reltio territory and we would tell you so. Plenty of organisations have both problems — the mistake is buying a tool for one and assuming it covers the other. Our post on duplicates covers the cleanup side honestly.
Frequently asked questions
What is the difference between data cleanup and master data creation?
Cleanup is retrospective: match existing duplicates across systems, apply survivorship, produce a golden record. Creation is prospective: someone needs a new supplier or cost centre, and a process decides what is captured, who validates and approves it, and which systems receive it. Cleanup fixes yesterday. Creation decides whether you repeat the cleanup in three years.
Why do MDM tools focus on matching and survivorship?
Because it generalises and it demos. Fuzzy matching works similarly across industries, so it can be built once and sold repeatedly, and a falling duplicate count is a compelling slide. Creation workflow encodes your specific organisation — who approves a credit limit — so it is unglamorous to build and unimpressive to demonstrate.
What happens if you only clean up master data?
Quality decays back to roughly where it started, because the process producing bad data is untouched. Deduplication removes thousands of duplicates and the same intake keeps producing them at the original rate. Companies that have run this exercise twice usually call it a data quality problem when it is a process problem.
What does good master data creation look like?
One front door for requests. Validation enforced at entry, not audited later. Domain values chosen from controlled lists. A duplicate check before creation rather than a match run after. Approval on the fields that carry risk, with a record of who approved what. And distribution to downstream systems from that single authoring point.
Fix the tap, not just the floor
Primentra is where master data gets made: controlled values, approval on the fields that matter, a full change log, and one place downstream systems read from. On your own SQL Server. Create a supplier end to end in the 60-day trial and time it.