No company starts out with a bad database. It starts with three hundred records written by the same person, all to the same standard, and at that point everything's fine.
What happens next is a slow decay that nobody decides on and nobody stops. And because it's slow, nobody notices until someone actually needs to use it for something serious.
How it goes wrong
There are four causes, and they usually all show up at once.
New people with new standards. Whoever joins was never handed a written rule — because there never was one — so they invent their own. After five people, there are five different ways of filling in the same field.
Free-text fields. Where you can type anything, people type anything. It's the single most productive source of inconsistency.
New tools with no rule for how they fit in. A new platform comes in, nobody decides how it connects with what's already there, and the same piece of data ends up existing in one more place.
Migrations. Every change of system loses something along the way: attachments, history, fields that had no equivalent on the other side and got dumped into a notes field.
The decay can be measured
And it's worth doing, because it turns a vague complaint into a number you can take to whoever decides the budget.
Pick three or four simple metrics and take them today: percentage of contacts with no valid email, percentage of records with the mandatory field left blank, number of likely duplicates, percentage of records nobody has touched in more than two years.
Take them again in six months. The difference between the two readings is the speed at which your database is decaying — and it's that speed, not the current state, that tells you whether it's worth doing anything about.
Cleaning it up once is money down the drain
The usual reaction is to commission a clean-up: someone spends three weeks removing duplicates, standardising addresses and fixing emails. Everything ends up spotless.
And then it goes right back to how it was, because nothing changed about the causes. Within a year and a half the database is back where it started, and now there's a feeling that it's already been tried and didn't work — which makes it far harder to get budget for a second attempt.
If there's only money for one thing, choose to stop junk getting in rather than clean up what's already there. A half-messy database where new records come in correctly improves on its own over time. A clean database with no rules at the door decays again, and faster, because in the meantime the volume has grown.
Upkeep: who, when, and by what rule
Data quality isn't a project with an end date. It's a small routine, and routines need three things clearly defined.
- Who. One named person per type of information. "The team" fixes nothing.
- When. One hour a month booked in the diary is worth more than one week a year.
- By what rule. Written down, on one page. How to write a company's name, what to do when a client who already exists shows up again, which fields can never be left blank.
And a warning from experience: if the written rule fights against whatever's quickest to do on screen, the rule always loses. It's better to change the screen.
Why this pays for itself
Because everything that comes afterwards rests on it. Reports you can trust, systems that connect without manual reconciliation, artificial intelligence that learns from your company's own records instead of learning from the mess in them.
It's the least visible work of all, and it's what decides whether everything else works — as we've written before about why nobody trusts the numbers.
Do you know the state of your database?
Tell us where your records live and what you use them for. We'll take the measurements with you and tell you what changes the rate of decay.