Most life sciences organizations have at least one: a data platform that hundreds of people rely on every day, built years ago, held together by nightly scripts, and understood in full by one person. Then that person moves on, the legacy source system is scheduled for retirement, and someone has to take it over without breaking what users depend on.
The instinct is to rewrite it. In our experience that is usually the riskiest option on the table. Legacy data pipeline modernization works better as a sequence: map the system, capture what it actually does in tests, put the schema under control, migrate one domain at a time, and run old and new side by side before anything is switched off. The platform stays live the whole way through.
What a bus factor of one looks like in a life sciences data platform
The "bus factor" (or truck factor) is the number of people who would have to leave before a system can no longer be maintained. It is lower than most teams assume. A study of popular open source applications on GitHub found that 46% had a truck factor of 1 and 28% a truck factor of 2. Those were public projects with visible code and contributor histories. Internal data platforms rarely have either.
In life sciences, the pattern tends to look like this:
One database carries both data entry and heavy analytics, so every reporting change risks the operational side.
Dozens or hundreds of materialized views are refreshed overnight by a script that only one person dares to touch.
Schema changes are made with hand-written SQL, with no migration history, so nobody can say with certainty what production looks like.
Development, test and production environments have drifted apart, and promotions are done by memory.
Records exist in the database but cannot be opened in the application, because a link, a date or a reference broke somewhere along the way.
The rules that explain why a value looks the way it does live in one person's head, not in code or documentation.
None of this is visible to the people using the dashboards. They see a platform that works. The risk sits underneath, and it surfaces at the worst moment: during an acquisition, a source system retirement, or a handover.
Why a big-bang rewrite is the riskiest way out
Rewriting from scratch feels clean. It promises a modern stack and an end to the fragile scripts. The problem is that a rewrite has to reproduce everything the old system does before it can replace it, including the behavior nobody documented.
The authors of Patterns of Legacy Displacement describe exactly this trap: even defining and agreeing current functionality becomes a huge effort and leads to a plan for a single "big bang" cut-over release. Martin Fowler's Strangler Fig Application makes the same point from the other side. Replacing a system piece by piece reduces risk, delivers value earlier, and lets the team learn before committing to the next step. Users do not stop needing the system while a rewrite is underway.
Life sciences adds one more constraint. Historical data is often the point of the platform. A pricing history, a reimbursement date or a past assessment record that disappears in a migration is not a cosmetic bug. It is a gap in the evidence people use to make decisions, and in regulated environments it is a gap someone will eventually have to explain.
A takeover sequence that keeps the platform live
The alternative is less dramatic and much safer. It treats the takeover as engineering work on a running system, not as a replacement project.
1. Map the system before changing it
Start with an inventory: entities, views, scheduled jobs, upstream sources, downstream consumers and owners. Which reports depend on which views? Which jobs must finish before others start? Which tables are written by the application and which only by scripts? This map becomes the shared reference for every decision that follows, and it removes the first layer of single-person dependency on its own.
2. Capture what the system actually does in tests
Before improving anything, record the current behavior in automated tests, including behavior that looks wrong. Michael Feathers called these characterization tests. They are not a statement that the current output is correct. They are a safety net that tells you immediately when a change alters something users rely on. Every later step depends on this net.
3. Put the schema under version control
Hand-written SQL applied directly to production is one of the clearest signs of a bus factor problem. Evolutionary database design, as described by Pramod Sadalage and Martin Fowler, treats every schema change as a migration script stored in version control next to the application code and applied the same way in every environment. Once that is in place, the database stops being something only one person can safely change.
4. Make environment drift visible
If development, test and production have drifted, you cannot trust any test result. Tooling that compares environments and generates reviewable promotion scripts turns drift from an unknown into a list of differences a reviewer can approve or reject.
5. Migrate domain by domain through one framework
Move data one business domain at a time, through a single, tested migration framework rather than a collection of one-off scripts. When a record cannot be resolved cleanly, carry it across in its original form and flag it. Do not drop it. A flagged record can be fixed later. A dropped record is usually discovered missing by a user, months after the fact.
6. Run old and new side by side before cutting over
Dual-run validation means both paths produce results for the same inputs and the differences are investigated before the old path is retired. The legacy displacement patterns call the supporting pieces a transitional architecture: components that exist only to make the move safe and are removed once it is complete. It costs some extra effort. It replaces a single high-stakes cutover with a series of small, checked steps.
A migrated record is not the same as a usable record
Many migrations are declared successful because row counts match. That is a weak test. The questions that matter are whether a user can open the record, whether its dates and references are intact, and whether the values resolve to the right entities in the new model.
Measure migration success by usability: records that can be opened, references that resolve, history that is complete, and exceptions that are flagged and countable. This is also where a takeover often finds value nobody expected, because data that was technically present but practically unreachable becomes usable again.
What this looked like beneath a live market access platform
We worked through this sequence on a data platform modernization beneath a global CRO's live market access platform. The platform had grown the way successful platforms do: one PostgreSQL database of around 171 entities serving both data entry and heavy analytics, 143 materialized views on a nightly refresh, schema changes applied by hand, a legacy source system due for retirement, and deep knowledge concentrated in very few people.
The work was delivered against the live platform, with no big-bang cutover:
8 legacy data domains migrated through one ETL framework.
100% of 38,571 pricing presentations (pack-level price records) migrated, up from 61.7%. Resolution reached 95.5%, and the remaining 1,721 were carried as free text and flagged, not dropped.
260,062 reimbursement dates recovered, and 12,129 clinical assessment and 5,587 reimbursement records made openable.
143 materialized views moved to an automated nightly refresh, backed by 142 test modules against 176 source modules.
Drift between development, UAT and production made detectable, with generated promotion SQL, and reporting pipelines made read-only by construction.
A five-phase architecture plan with dual-run validation for the steps still ahead.
Sometimes a component cannot be moved at all. In a multi-cloud infrastructure consolidation into Azure after an acquisition, systems that could not be migrated were rebuilt instead, with zero data loss and no user-facing disruption, and deployments went from days to under an hour. Knowing when to rebuild a piece, rather than the whole, is part of the same discipline. For the integration questions that follow a deal, see what breaks after the deal closes.
Warning signs you are closer to this problem than you think
Nobody can explain, without asking one specific person, what happens if the nightly refresh fails.
Schema changes are not in version control.
Test and production give different answers to the same query, and nobody is sure why.
A source system is scheduled for retirement and there is no written migration plan.
Users report records they can see in a list but cannot open.
The most experienced engineer on the platform has not taken a real holiday in a year.
Two or more of these usually means the platform is one resignation away from a crisis. It also tends to explain why digital products in this sector lose momentum after launch: once the original builders move on, every change gets slower and riskier.
A practical starting point for a legacy data takeover
You do not need to commit to a modernization program to reduce the risk. A landscape assessment maps the data sources, systems, dependencies and owners, and shows where the single points of failure are. From there, one defined workstream (one domain migrated, the schema put under version control, or the test net built) proves the approach on your own system before anything larger is planned.
This is the kind of work our data engineering team for life sciences does: taking over live systems, keeping them running, and handing them back documented and testable so your own people can own them. Where the takeover reaches infrastructure, the same approach extends to cloud migration and DevOps.
The goal is a platform more than one person can run
A successful takeover is not measured by how modern the new stack looks. It is measured by whether the platform kept running, whether every record survived, and whether the next engineer can change it safely without asking the last one. Treat the legacy system as a source of truth to be understood before it is replaced, and the rewrite you were afraid of often turns out to be unnecessary. For the broader case on why the operational data layer, not the tooling on top, decides what a platform can do, read why the data operations layer is the real constraint.