Every time you open Netflix, join a video call, or pay through your banking app, there is a high chance that platform changed while you were sleeping.
And you probably never noticed.
Most people think stability means not touching the systems. After more than ten years in telecommunications, infrastructure, and critical platforms, the reality is usually exactly the opposite.
Stability exists because those systems change constantly. Not because something broke. But because technology never stops evolving.
New features, performance improvements, security updates, hardware replacements, automation, migrations… all of that requires taking changes into production. And most of it happens overnight, when the impact on users is as low as possible.
The real challenge comes before the console
Over the years it becomes clear that the hardest part of this work is not executing a change. It is still fun to do it once in a while, but the real challenge comes long before opening a console.
It is about making sure everything arrives prepared for that moment.
In practice, a large part of the work is aligning teams across different countries, reviewing the pipeline for the coming weeks, understanding which projects are coming, identifying dependencies, raising risks before they appear, and making sure everyone shares the same context before going into production.
Those conversations often start weeks before the maintenance window. There are meetings with engineering, operations, vendors, architecture, customers, and management. There are competing priorities. There are dates that look simple on paper, but hide weeks of preparation.
Saying yes to everything does not always help
And one of the biggest things I have learned is that saying “yes” to everything does not always help the business.
Sometimes the greatest value you can add is raising your hand in time. Explaining why a risk exists. Showing the impact. Proposing an alternative. Adjusting expectations before the problem reaches production.
At first those conversations can be uncomfortable. Later you realize they build far more trust than promising something you know is not realistic.
It also matters that the team’s work is visible. If an effort is not planned, broken into tasks, estimated, and prioritized, it becomes very hard to show the real workload, justify resources, or make objective decisions.
That is why it is worth turning conversations into planning. Not to fill Jira. But because it is one of the best ways to protect the team and make decisions based on data instead of perceptions.
When the plan is no longer enough
This is not anecdotal. In critical infrastructure, change is the normal state.
British Telecom has talked about roughly 11,000 network changes per week. Airtel has reported on the order of 3.1 million changes per month across its network. And the same happens outside telecommunications: banks, streaming platforms, cloud providers, and digital services deploy updates, patches, and migrations almost every day. Much of that happens outside visible hours, when the impact on users is lower.
That is why so many platforms look stable: not because nobody touches them, but because someone is touching them carefully.
A few weeks ago we lived one of those typical infrastructure nights.
It was an upgrade we had already executed before. The runbook had been reviewed and everything seemed under control. However, because of a completely human error, we skipped a critical step before starting.
The technology did not fail. The procedure existed. The script existed. A step simply was not executed.
When we realized it, the change was already underway. We had to adapt: get into servers, apply the pending configuration, validate connectivity, and continue. Work that was meant to finish around five ended closer to eight. We opened an incident, extended the window, kept the teams informed, and waited for full validation.
In the end the upgrade was successful. Users never noticed what happened behind the scenes. And that was exactly the goal.
At another scale, the same thing happens. In February 2026, Cloudflare published a postmortem about an internal change that, because of a bug in a cleanup task, withdrew network prefixes by mistake. It was not an attack. It was a poorly bounded change, with global impact and hours of recovery.
The lesson is the same at a different scale. Modern infrastructure does not depend only on scripts and runbooks. It depends on people’s judgment, on preparation, and on the ability to decide when the plan is no longer enough.
Technology, judgment, and people
AI is also changing this world. Today it is possible to generate scripts, automate repetitive tasks, analyze configurations, and even build internal tools that help manage complex projects better.
Productivity is very different from a few years ago.
But there is still something that remains deeply human.
Judgment. Knowing when to move forward. When to stop a change. When to ask for help. When to trust the team. And when to take responsibility for a difficult decision.
Maybe that is why this work still makes sense. Because many of those changes that look small end up enabling enormous projects. A new service. A massive migration. An entirely new platform.
From the outside they look like isolated tasks. From the inside you know everything is connected.
The next time you open a streaming platform, join a video call, or simply browse the Internet, you may be using infrastructure that changed while you were sleeping.
If you never found out…
It is because many people did their job well.
And that, curiously, is still the best sign that a change was truly successful.
Related reading:
- Project Management: Make Things Happen for Real
- When Jira Is Not Enough
- From VTR Chile to the team in Europe
✍️ Claudio from ViaMind
“Dare to imagine, create and transform.”