As Rolando Bonaccorsi, Director of Operations at Vert Analytics, points out, cloud computing allows companies to maintain systems, applications, and data on infrastructure accessed over the internet, reducing their dependence on in-house servers and increasing operational flexibility. However, even widely used platforms are subject to failures that can disrupt services and compromise essential operations.
Why Can a Cloud Failure Affect So Many Services?
A seemingly independent application may rely on several external resources to function. Databases, authentication, storage, and communication between systems often depend on shared services. When one of these components experiences problems, different applications can be affected simultaneously, even if their code continues to function properly. This interdependence can trigger cascading failures, expanding the impact of an outage that was initially limited to a single service.
In addition, concentrating resources within the same infrastructure region can amplify the consequences of an outage. If critical systems rely exclusively on that location, a regional failure can compromise access to information and prevent operational processes from being carried out. Resource distribution must take these dependencies into account. Keeping essential components in different regions can reduce exposure to localized incidents, provided that communication and recovery mechanisms are also prepared for this configuration.
Another factor is the difficulty of quickly identifying the source of the problem. An unavailable application may display similar symptoms when facing network, authentication, or database failures. Therefore, as Rolando Bonaccorsi emphasizes, understanding the architecture and mapping the relationships between components is essential to guide the initial response. Monitoring tools and operational logs help identify which services have been affected and distinguish the root cause from problems resulting from the outage.
What Needs to Keep Running During an Outage?
Not all services have the same level of importance when it comes to business continuity. Systems responsible for financial transactions, customer service, communications, or the processing of critical information may require priority recovery. Certain administrative activities, on the other hand, may tolerate longer periods of downtime without immediately compromising operations.
From the perspective of Rolando Bonaccorsi, whose professional background includes technology operations and delivery, this distinction highlights the need to establish priorities before incidents occur. Planning should identify which processes support essential operations and which technology resources are indispensable to keeping them running.
How Can Recovery Be Prepared Before a Failure Occurs?
Planning begins with defining two parameters: the maximum acceptable amount of time to restore a service, known as the Recovery Time Objective (RTO), and the amount of data loss an organization is willing to tolerate, known as the Recovery Point Objective (RPO). These objectives guide the selection of technical solutions and help determine the level of investment required to protect each operation. The stricter these limits are, the greater the need tends to be for redundant resources, rapid recovery mechanisms, and procedures capable of reducing the impact of outages.
Available alternatives include backups, data replication, and secondary environments prepared to take over specific workloads. However, maintaining resources in another region does not automatically guarantee continuity. It is necessary to verify whether authentication, connections, permissions, and other dependencies will also be available during recovery. Compatibility between environments and keeping replicated information up to date are equally important to prevent inconsistencies when services are restored.
Rolando Bonaccorsi’s experience in IT operations management also highlights the importance of having predefined procedures in place. Teams need to understand their responsibilities, the criteria for activating alternative environments, and the appropriate recovery sequence. Without this preparation, decisions made during a crisis can extend the period of downtime. Regular testing and simulations help verify whether procedures remain effective and whether teams can execute them within the established time frames.
