New! Listen to Concept to Cloud - Real stories from the trenches of software engineering

Service · Legacy Modernisation

Data platform modernization, one workload at a time.

The same gated method as our application modernization work, pointed at the data estate: inventory the pipelines and reports, give every workload a disposition, dual-run one migration until the numbers reconcile, then scale by priced milestones. No big-bang replatform, ever.

What is data platform modernization?

Data platform modernization is moving an estate of pipelines, databases, and reporting workloads from aging infrastructure onto modern, governed foundations: managed services where they earn their keep, versioned and tested transformations, lineage you can trace, and a serving layer both humans and AI systems can trust. The unit of work is the workload, not the platform: each pipeline and report gets its own decision, and the migration happens one proven slice at a time.

It sits on the seam between our legacy modernization and data engineering practices, and it is usually what a stalled AI initiative is actually waiting on: the argument is in why your AI project is actually a data project.

The diagnosis

How do you know it is time?

Data platforms rarely fail loudly. They accumulate symptoms until the estate is somebody's full-time apology. The usual six:

  • The overnight batch finishes at 11am, and nobody can say which half of it still matters
  • Spreadsheet exports are load-bearing integration between systems
  • One person knows how the ETL works, and they have opinions about holidays
  • The reporting database is also the production database, and both are slow
  • Every AI initiative stalls at "first we need to sort out the data"
  • The cloud bill grows faster than usage, and nobody can attribute it to workloads

The method

How do you modernize a data platform?

In four gated steps: inventory the workloads, disposition each one with the 7 R's, prove the migration pattern on a single pipeline with a dual-run, then scale into a funded milestone plan. It is the data-estate application of our application modernization roadmap, with one addition the data version demands: old and new run side by side until the numbers reconcile, because in data work "it deployed" proves nothing.

What the dual-run produces

RunDelta vs legacyWhat happened
Day 1-0.53%Timezone truncation found in the legacy rollup; matched, then documented
Day 9-0.02%One source mapping corrected in the new path
Day 160.00%First clean agreement; the consecutive-run counter starts
Day 370.00% × 21 runsConsumers moved; the legacy job switched off, unremarked
An illustrative reconciliation log for one workload: the numbers are invented, the format is the real deliverable. This table is the whole argument. After three weeks of zeros on real runs, cutover requires no faith from anyone.

Step 1

Typically 2 weeks

Inventory the workloads

Know every pipeline, database, and report, and who actually uses each.

  • List every pipeline, scheduled job, database, and report with an owner and a consumer
  • Check the access logs: reports nobody has opened in a quarter go on the retire list
  • Map the dependency chain from source system to the numbers leadership reads
  • Record freshness, run time, and failure history for every scheduled job

Exit gate

Every workload has an owner, a consumer count, and a known dependency chain.

Step 2

Typically 2 weeks

Disposition every workload

The 7 R's, translated to a data estate.

  • Retire the dead reports and the pipelines that feed only them
  • Retain what works; rehost or replatform databases onto managed services, or relocate virtualized hosts wholesale when a data-center exit sets the clock
  • Repurchase where a product does it better (BI tools, connectors, orchestration)
  • Refactor into a modeled warehouse layer only the workloads that earn it

Exit gate

A disposition matrix for the estate, signed off, with the first migration candidate chosen.

Step 3

4-8 weeks

Prove it on one workload

Migrate one pipeline end to end, running old and new side by side.

  • Build the new path with versioned, tested transformations and lineage from day one
  • Dual-run old and new, reconciling outputs until the numbers agree for real users
  • Cut consumers over gradually; the old pipeline dies by disuse, never by decree
  • Publish the before-and-after: run time, failure rate, cost per run

Exit gate

One workload live on the new platform with reconciled numbers and its old path unplugged or atrophying.

Step 4

1-2 weeks

Scale into the funded plan

Turn the proven pattern into a milestone-priced migration of the rest.

  • Sequence the remaining workloads by dependency and value
  • Price each migration phase with acceptance criteria and stop-points
  • State what is explicitly staying put and why
  • Present with the first workload's numbers as the evidence

Exit gate

The next phase funded in writing, with the first workload's numbers doing the arguing.

Frequently asked

Questions we get

What is the difference between data modernization and data platform modernization?

Data modernization is the broad program: getting an organization's data estate onto modern foundations. Data platform modernization is the engineering core of it: replacing the pipelines, databases, and serving layers the data actually runs on. In practice they travel together, and the platform work is where the budget and the risk live, which is why this page is about the platform.

What are the 5 layers of a data platform?

Ingestion, storage, transformation, serving, and governance. Modernization efforts usually obsess over storage (the warehouse or lakehouse choice) and neglect transformation and governance, which is where estates actually rot. Our data engineering practice assesses against all five.

How long does data platform modernization take?

The first workload should be live on the new platform inside two months; the four steps above total roughly 9-14 weeks to a funded plan with that migration shipped as proof. A full estate takes quarters, not weeks, and the honest number for yours comes out of the first two steps. Anyone quoting a total before the inventory exists is guessing.

Should we modernize the data platform before or after the applications?

Dependency order decides, and it usually says data first: applications consume the data layer, so a modernized application on a rotten data platform inherits the rot. The two programs share one method here; the application side is covered in our application modernization roadmap, and the disposition matrices from both should sit on the same table.

Does the warehouse get rebuilt as part of this?

Sometimes. The disposition step decides: a warehouse with a broken model gets remodeled subject area by subject area, while a well-modeled but expensive one usually needs replatforming rather than rebuilding. The warehouse-specific version of this work, design, build, and rescue, lives on our data warehouse consulting page.

Start with steps 1 and 2.

The Technical Assessment is steps 1 and 2 run full-time instead of alongside the day job, which is how two typical-2-week steps fit into a one-to-two week engagement. Fixed price agreed before we start. You leave with the estate scored and the first migration scoped and priced; the milestone plan for the remainder follows the dual-run's evidence, exactly as step 4 says. Every artifact is usable with us or without us.