Data cleanup and migration
Moving data is straightforward. Moving it correctly — deduplicated, reconciled, with the judgement calls recorded and reversible — is the part that decides whether the new system is trusted or quietly worked around.
The migration that poisons the new system
Bad data carried across doesn't stay contained; it discredits the system it lands in. Staff spot a handful of wrong records, conclude the new tool cannot be relied on, and go back to the spreadsheet they trust. The build was fine. The migration is what lost them.
You’ll recognise this if
- You are replacing a system and have to bring history with you
- The same customer or product exists several times under variant names
- Records live across spreadsheets that have diverged from each other
- A previous migration left data nobody fully trusts
What we actually deliver
Deduplication with the rules written down
Matching that handles the realistic cases — punctuation, abbreviations, transposition, changed addresses — with the merge rules documented and the merges reversible.
Standardisation and validation
Dates, addresses, phone numbers, currencies, and identifiers brought to one consistent form, with anything that fails validation quarantined for review rather than silently dropped.
A repeatable, rehearsed migration
The migration is code, run repeatedly against a copy until it is clean. Cutover day executes something already proven rather than something attempted for the first time.
Reconciliation you can show someone
Counts and totals proving what came across, what was merged, and what was held back — evidence for finance, for audit, and for the staff who need to believe the new system.
The way we approach it
Decide who is authoritative first
Where three systems hold the same customer, one has to win, field by field. This is a business decision more than a technical one, and we get it agreed before writing any transformation.
Never destroy the source
The original data is preserved untouched and every transformation is re-runnable. If a rule turns out to be wrong three weeks after cutover, it can be corrected rather than mourned.
Let the people who know the data review it
We surface the ambiguous cases — the near-duplicates, the contradictions — to the people who can actually adjudicate them, in a form that makes reviewing hundreds of them tolerable.
What changes
- One trustworthy record per customer, product, or entity
- A migration that has been rehearsed before it matters
- Reconciliation evidence that satisfies finance and audit
- Staff who trust the new system enough to abandon the old one
Asked often enough to answer here
How much downtime does cutover need?
It depends on volume and whether a phased approach is possible, but rehearsing the migration repeatedly means we can tell you the number in advance rather than discovering it live. For many systems it is a matter of hours.
What happens to records that can't be cleaned?
They are quarantined with the reason attached, never silently dropped. You get a list and decide — some will be worth manual attention, and some will be genuinely dead records worth retiring.
Can you do this without replacing our system?
Yes. Cleanup within an existing system is a common engagement in its own right, particularly ahead of an analytics or AI project that needs the data to be dependable first.
Related
Wherever you’re starting from, let’s figure out the next step.
Tell us what you’re building — we’ll tell you honestly whether we’re the right team for it.