Back to blog
PrimentraPrimentra
·July 25, 2026·8 min read

Master data vs reference data: two different problems people keep solving with one tool

Home/Blog/Master data vs reference data: two different problems people keep solving with one tool
REFERENCE DATA
NLNetherlands
DEGermany
FRFrance
PTPortugal
249 rows · changes yearly · ISO defines it · no duplicates possible
MASTER DATA
ABC Corp
A.B.C. Corporationdup?
ABC Company Inc.dup?
ABC Corp.dup?
84,000 rows · changes hourly · you create it · duplicates are the whole problem
Same governance programme. Completely different problems.

Reference data is the controlled code lists your systems look values up in: country codes, currencies, units of measure, status flags, GL account categories. Master data is the business entities you create and operate on: customers, products, suppliers, employees, cost centres. Reference data is small, changes rarely, and is often defined by somebody outside your company. Master data is large, changes constantly, is created by your own staff, and is mostly a duplicate problem. They get filed under one heading and governed as one thing, and that is where the trouble starts.

The relationship between them is worth stating plainly, because it is the part that gets lost: reference data is usually referenced by master data. A customer record has a country attribute, and that attribute points at the country list. One is the vocabulary; the other is what you write with it.

Where they actually differ

Reference dataMaster data
SizeTens to hundreds of rowsThousands to millions
Change rateRarely — a change is an eventContinuously, by many people
Who defines itOften an external standard (ISO, UNSPSC)Your business, record by record
IdentityThe code is the identity. No ambiguityAmbiguous. Matching is the core problem
Typical ownerNobody, which is the problemA business function — procurement, finance, HR
What it needsVersion control and reliable distributionStewardship, approval, duplicate prevention
How it failsLists drift apart. Joins and imports breakDuplicates and conflicting values. Wrong decisions

The row that matters most is identity

There is exactly one Portugal, and its code is PT. You cannot accidentally create a second one, because the code is the identity and the list is short enough to see in one screen. Reference data has no identity resolution problem at all.

Master data is almost entirely an identity problem. Is ABC Corp the same company as A.B.C. Corporation? Nothing in the data answers that. This single difference is why the two need different tooling: one needs a way to agree on a list, the other needs a way to decide whether two records are one thing.

What conflating them costs

Buying matching for a list problem. A team decides reference data is out of control and evaluates MDM platforms built around fuzzy matching and survivorship. None of that helps. Their forty-nine currencies are not ambiguous; they are just inconsistent across four systems. What they needed was one owner and a distribution mechanism, at a fraction of the cost.

Governing customers like a code list. The reverse is more common and more expensive. A spreadsheet genuinely is adequate for a currency list that changes once a decade. Applied to eighty thousand customer records edited by fifteen people, the same approach produces exactly the duplicate estate that funds the consulting industry.

Leaving reference data unowned. The most common failure of all, and it comes from reference data looking too small to matter. Nobody is assigned to it, so each system keeps its own copy, and the copies drift. One system uses NL, another NLD, a third still carries a status value the others retired two years ago. Then the integrations start dropping rows, quietly. We went into that failure mode in the part of your MDM strategy that breaks first.

Govern them differently, in the same place

The practical answer is not two systems. It is one system that understands the two behave differently. Reference data wants a named owner, a change history, and one authoritative copy every system reads — the emphasis is on distribution. Master data wants stewardship, validation at entry, approval on risky fields, and duplicate prevention — the emphasis is on the process of creating a record.

In Primentra both live as entities in the same model, and the link between them is a domain-based attribute: the Customer entity's country attribute points at the Country entity, so a steward picks from the governed list rather than typing. That is what makes the vocabulary actually get used — not a policy document, but a dropdown that only contains valid values. If you are working out which of your data is which, what is reference data covers the smaller half, and master vs transactional data covers the other axis people confuse.

Frequently asked questions

What is the difference between master data and reference data?

Reference data is controlled code lists your systems look values up in — countries, currencies, units of measure, status flags. It is small, changes rarely, and is often externally defined. Master data is the business entities you create and operate on — customers, products, suppliers, cost centres. It is large, changes constantly, is internally created, and suffers from duplicates. Reference data is usually referenced by master data.

Is reference data a type of master data?

Many frameworks treat it as a subset, which is harmless on an org chart. Operationally they differ enough that managing them identically causes problems: reference data has no identity resolution problem because the code is the identity, while master data is mostly an identity problem.

Why does the distinction matter in practice?

It determines what you buy and how you govern. Reference data needs an owner, version control, and reliable distribution. Master data needs stewardship, approval workflow, and duplicate prevention. Conflating them leads teams to buy matching engines for list problems, or to manage customers with the informality that genuinely does work for a currency list.

What breaks when reference data is not governed?

Joins and imports, usually first. Two systems carry slightly different versions of the same list — NL versus NLD, or a retired status value one system still uses — and integrations silently drop or reject rows. Reference data is small enough that nobody is assigned to own it, which is precisely why it drifts.

One model, both kinds of data

Primentra holds your code lists and your business entities in the same model on your SQL Server, and links them with domain attributes so stewards pick valid values instead of typing them. Load a country list and a customer entity into the 60-day trial and see the dropdown do the governing.

Start free 60-day trial →Reference data management →Cleanup vs creation →

More from the blog

Payment terms master data: the Net 30 that was really Net 448 min readCurrency master data: the stale exchange rate that cost €40,0008 min readProduct classification master data: why your biggest spend category is Miscellaneous8 min read

Ready to migrate from Microsoft MDS?

Join the waitlist and be the first to try Primentra. All features included.

Download Free TrialTry DemoCompare MDM tools
Master data vs reference data: two different problems people keep solving with one tool | Primentra