Migrations always look fairly tidy in the planning. There is a current platform, a new platform, a date, and a series of activities that should take us from one point to the other.
Then you actually start working on the migration, and the dependencies show up, along with the teams, the change windows, the validations, the rollbacks and all those small things that did not fit in the first line of the roadmap.
Over the years there is something I have seen repeat itself quite a lot: problems tend to show up in the details. Not necessarily because someone did their job badly, but because the more critical and the older a platform is, the more things exist around it. And many of them only become visible when you start moving them.
The scale changes. Some rules not so much.
One of my first experiences with migrations was in Chile, working on network infrastructure for medical centers. We had to renew fairly old networks. That meant preparing switches and routers in advance, creating VLANs, leaving the configurations ready and then physically going to the medical center to make the change.
The windows were usually in the middle of the night. I remember being there, taking down old infrastructure, bringing up the new network and coordinating with the provider to activate connectivity to the corporate MPLS. Until everything was working again, the job was not done.
Today projects have a different scale, but several rules are still the same: prepare in advance, know what has to happen, have the right people available, execute in a clear order and validate before moving on.
A migration almost never affects only the platform you are changing
Over time, projects grew considerably in size and complexity. In a video platform migration in Switzerland, for example, the change was not limited to hardware. We also had to move local components, update packager software, review profiles, adapt configurations and make sure service distribution kept working correctly.
Each of those pieces had its own dependencies. And each one had to change at the right moment. That kind of project makes something very clear that is not always visible from the outside: a platform almost never lives alone.
You can change the hardware correctly and still have a problem further down the line. You can have the software working and fail on distribution. You can have the services available and discover a local configuration nobody had considered. That is why understanding dependencies before production ends up being one of the most important parts of the job.
A lot should already have happened before production
My work is usually quite close to deployment and to the arrival in production. But when the change window starts, much of the important work should already be done.
First, we try to reproduce as much as possible in pre-production. We validate functionality, run the steps, review integrations and confirm what should happen after each activity. That is where the runbook comes from.
And a good runbook should not be just a list of commands. It has to make clear what we do, in what order, who does it, who validates, what we expect to see, what we do if something does not happen as expected and how we go back.
In a big window you can have many people connected at the same time. That does not guarantee coordination. If responsibilities are not clear beforehand, having twenty people on a call does not solve much.
Production is always a different animal
You can test a lot. And you should. But pre-production is never exactly production. In production you get real users, real load, more services, configurations that have existed for years and combinations you cannot always fully reproduce beforehand.
On top of that, that is where the business is: the customers, the KPIs, the operation and the economic impact of a problem.
That is why passing every test does not mean we can assume everything will behave exactly the same. It means we have reduced the risk considerably. Production is still ahead.
That is why I try to start small
One of the things that has made the most sense to me in big migrations is avoiding moving too much at once. If we can start with ten channels, we start with ten. If we can start with one package, we start with one package.
We migrate, validate, measure, listen to operations and observe. If a problem appears, the impact is contained. If everything works, the next batch starts with much more information.
That can look slower in the planning. But finding a problem after moving ten services is very different from finding it after moving five countries.
Migrating in batches also means leaving time to learn
In one of the projects I am currently working on, we are migrating a DRM platform. The change affects linear and non-linear content, replay, VOD, back office integrations and local components.
There is no button that says “migrate DRM”. There are several changes. First, prepare. Then move one part, keep the ability to go back for a while, observe and only then continue.
That is why after certain changes we leave stability periods. Some problems show up immediately. Others need traffic. Others appear on a weekend or with a specific feature.
The weeks when apparently nothing happens are also part of the migration. We are not waiting for the sake of it. We are waiting for evidence.
The rollback has to really exist
While we are still learning, it is important to keep a clear way to go back whenever it is technically possible. Operations has to know what changed, what needs to be reverted, in what order and who validates afterwards.
It is not enough to put a line in the plan that says “Rollback available”. It has to remain possible to execute it. Because as the migration progresses, there may come a point where going back stops being simple.
That moment should be a conscious decision. Not something we discover when we already need to do it.
And while you migrate, the rest of the work goes on
A major migration is rarely the only thing the team is doing. There are still upgrades, maintenance, other deliveries, production issues and requests from different operations.
So planning starts to become critical. You do not only need to know whether you can technically execute a change. You also need to know whether people are available, whether another activity depends on the same platform, whether there is an operational window and what happens to the rest of the plan when a date moves. In large projects, a small deviation at the beginning can end up shifting many things at the end.
And that is an important part of project management. Not making the project always look under control, but making visible what is confirmed, what still carries uncertainty and what could move if any of the assumptions change.
Deadlines matter. A lot. But bringing a date forward does not always mean speeding up a migration. Sometimes it simply means taking on more risk.
Today we also have better tools
We are starting to use AI-based tools to help us with timelines, dependencies, diagrams and plans that are starting to have too many pieces. Not to make decisions for us. Simply to get better visibility and reduce part of the manual work around planning.
In the end, I keep coming back to the details
Big things usually get a lot of attention. Everyone knows we are changing a platform. Everyone knows there is an important date.
What ends up worrying me more are the small things around it: a different configuration in one country, a backend call that changes, a runbook step that has to happen before another, a validation everyone thought another team was doing, an old dependency that was not documented. That is where many of the surprises show up.
That is why a migration needs architecture, technical knowledge and planning. But it also needs a lot of attention to detail. Not because we can anticipate absolutely everything, that is impossible, but because the more assumptions we manage to confirm before production, the fewer surprises we leave for the change window.
A migration is not one change night
Maybe this is the simplest way I see these projects today. The migration does not start when we open the window. And it does not end when we close the change request.
It starts much earlier: understanding dependencies, testing, preparing, coordinating teams and building the runbook. Then production arrives. We move one part, measure, wait, learn, correct and scale up.
The goal is not to make sure a problem never appears. In large systems that would be unrealistic. The goal is that, when it does appear, the impact is manageable, we have enough information to understand it and we still keep a way to go back.
For me, that is a big part of the real work behind a complex migration.
Related reading:
✍️ Claudio from ViaMind
“Dare to imagine, create and transform.”