When Three Different Vendors Each Say the Outage Isn’t Their Fault

Three IT team members arguing during an outage, illustrating vendor blame in multi-vendor outages

Summary

Most companies can’t say how long their last major outage really lasted, not what the status page eventually reported, but the true duration from start to finish. When multiple vendors are involved, this becomes especially hard to pin down, because each vendor’s timeline is technically true from its own perspective. However, none of these timelines, on their own, add up to the full story.

How long was your last major IT outage really? Not what was ultimately stated on the status page, but the actual duration of the event from start to finish. In cases where multiple vendors are involved, most companies have no idea how to answer this question, because each vendor’s timeline is technically true from its own perspective. However, none of these timelines offer the whole story.

This is slowly becoming one of the more neglected problems of IT service management (ITSM), but certainly not because outages are rare. Outages have become increasingly common and often involve multiple vendors.

Why multi-vendor IT outages are becoming the norm

The modern IT ecosystem is based on layers of dependency. An enterprise is likely to run its own software solutions based on cloud infrastructure, which then uses an identity management service that uses different SaaS solutions and their respective vendors. If any part of this chain breaks down, its effects are unlikely to be confined to that specific solution or organization.

It’s a structural characteristic of modern-day IT environments rather than some rare exception. According to Zylo’s SaaS Management Index, enterprises currently have an average of 305 SaaS applications. Simultaneously, the IT infrastructure underpinning these apps is a lot less diverse than the variety of applications might suggest. In Q1 of 2026, AWS, Microsoft Azure, and Google Cloud controlled 63% of the entire global cloud infrastructure market, according to information from Synergy Research Group provided by Statista.

The reality is that a lot of applications that don’t seem to be related to one another are built on top of the same limited set of underlying platforms. With the growing number of external systems involved in companies’ operations, the likelihood of an outage affecting multiple companies naturally increases. A single failure in the system may affect multiple companies that were not responsible for it and didn’t even know about it.

Why vendor incident timelines rarely match

When an outage involves more than one vendor, it is very common for each vendor to have its own timeline for the same event. This does not necessarily mean that anyone is lying about the actual timeline, but rather that each vendor is basing its timeline on its own internal metrics.

Service level agreements (SLAs) generally define uptime and severity in such a way that leaves a lot to interpretation. Whether a partial outage is considered a down state, or a certain amount of time in a suboptimal state of performance should be considered downtime, is not clearly defined. Additionally, vendors use internal definitions of incident severity, which affect the metrics for their response and resolution times. Finally, vendor uptime statistics are compiled from the vendor’s internal data.

None of this necessarily involves anything wrong. Every vendor bases its metrics on what it sees within its own network.

Vendors and outages

The business cost of unclear incident accountability

Apart from the actual downtime, the cost of establishing whose system failed also comes into play. According to ITIC’s Hourly Cost of Downtime Report, 97% of large enterprises (those with over 1,000 people) agree that downtime of one hour costs more than $100,000, the average cost of an hour of downtime is higher than $300,000 for over 90% of mid-size and large enterprises, and 41% of enterprises incur hourly downtime costs in the range of $1 million and more than $5 million. Each hour spent sorting out vendor accounts takes away time from addressing the actual cause of the problem or informing employees and customers about it.

In case of a repeat incident, it is difficult to determine whether it is the same underlying problem or an altogether new incident because there is no independent, documented account of the previous incident.

Why an independent incident record matters more than vendor agreement

Coordinating multiple external vendors’ accounts into a single version of events in a case like this is unrealistic. Their tools and definitions of events will be different.

The only aspect an organization can affect is whether it has an independent log of events it reviewed during the incident. This includes logging events as they happen, not trying to recreate them later using third-party information. Incident categorization should be done consistently rather than assigning random labels to events. Additionally, it is important to maintain an internal timeline of incidents without relying on external parties.

This is primarily about discipline in using the current tools, rather than anything else. It includes structured incident categorization, accurate timestamps, and reporting infrastructure that can recreate the actual sequence of events. Modern ITSM platforms provide the required structure to do this: They can define categories and subcategories, and they have reporting and analytics infrastructure that can filter and recreate the timeline based on the actual logs.

Considering internal causes, not just external ones

For a balanced perspective on incident accountability, one should also consider the organization’s internal procedures. Human error accounts for about 66–80% of all downtime incidents according to Uptime Institute data cited in FirstPassLab, where the two biggest contributors to such problems are employee failure to follow the established procedure (47%) and procedure flaw or uncertainty (40%). According to the same report, only about 3% of organizations manage to identify their own mistake in advance before the outage occurs.

This is why it makes sense to highlight that the primary objective of the independently maintained incident record is not to prove the fault of external vendors, but to document the incident regardless of its cause.

How log data supports incident verification

Incident tickets document the report and its resolution, but they rarely detail the issue’s root cause at a more technical level. Sometimes, a seemingly straightforward technical problem, whether it originates in the organization’s infrastructure or the vendor’s system, has a security angle. For example, stolen credentials used to initiate some actions or network traffic following a pattern known to be malicious.

It is information that cannot be uncovered through incident tickets, but can be uncovered through analysis of logs and events. A Security Information and Event Management (SIEM) solution continuously gathers logs from firewalls, servers, and other systems, comparing that activity with threat intelligence feeds – databases of malicious IP addresses, domains, and attacks. If there is a match between the incident’s activity and any of those indicators, the match is identified and confirmed. If there is no such match, then a security-related cause is unlikely. In either case, the answer is provided by information available within the organization’s internal environment.

The key takeaway for ITSM management teams

It seems multi-vendor outages will only become more frequent as IT environments continue to get more intertwined, and there is nothing an organization can do to force different vendors to reach an agreement on what happened during an incident. However, this is not the extent of an organization’s power; it is the beginning.

The one thing that is entirely under the control of an organization is how prepared its own IT service desk is in anticipation of the next incident; this includes categorizations that remain consistent from the very first ticket, timestamping throughout the course of an event rather than retroactively, and the capability to produce a report that allows for a consistent, defensible chronology in case it becomes necessary. This is not a gap that can be solved by acquiring additional technology. Instead, this problem must be solved using the capabilities already available in many ITSM solutions, and it must be addressed proactively rather than reactively.

Vendors FAQs

Why are multi-vendor outages becoming more common?

Modern IT runs on layers of dependency. Enterprises use an average of 305 SaaS applications, according to Zylo’s SaaS Management Index, but most of that software sits on a much smaller set of underlying platforms. AWS, Microsoft Azure, and Google Cloud controlled 63% of the global cloud infrastructure market in Q1 2026, per Synergy Research Group data cited by Statista, so a single failure can ripple through companies that had no direct involvement.

Why don’t vendor timelines match after a shared outage?

Not because anyone is lying. Each vendor bases its timeline on its own internal metrics, and SLAs often leave terms like uptime and severity open to interpretation, so what one vendor calls “down” another may not.

What does unclear incident accountability cost a business?

According to ITIC’s Hourly Cost of Downtime Report, 97% of large enterprises say an hour of downtime costs over $100,000, and 41% put their hourly cost between $1 million and $5 million or more. Every hour spent sorting out whose system failed is an hour not spent fixing the problem or informing employees and customers.

Can an organization force vendors to agree on one version of events?

No. Coordinating multiple vendors’ tools, definitions, and accounts into a single timeline isn’t realistic, according to the article.

What can an organization actually control during a multi-vendor outage?

Its own independent incident log: logging events as they happen rather than reconstructing them afterward, consistent categorization of incidents, accurate timestamps, and a reporting infrastructure that can recreate the real sequence of events. Modern ITSM platforms already support this.

Are multi-vendor outages always the vendors’ fault?

Not necessarily. Human error accounts for an estimated 66 to 80% of downtime incidents, per Uptime Institute data cited in FirstPassLab, with employees failing to follow procedure (47%) and flawed or unclear procedure (40%) as the two biggest contributors. Only about 3% of organizations catch their own mistake before an outage happens.

How does log data help verify what actually caused an incident?

Incident tickets record what happened and how it was resolved, but rarely the technical root cause. A Security Information and Event Management (SIEM) system continuously compares log activity from firewalls and servers against threat intelligence feeds, which can confirm or rule out a security angle that a ticket alone wouldn’t surface.

Judin Joan Soundarya S
Judin Joan Soundarya S
Product Marketing Specialist at ManageEngine
Judin is a cybersecurity specialist at ManageEngine, a division of Zoho Corporation. Her work spans blogs, case studies, and campaign content that break down complex security ideas into crisp, accessible insights. She translates customer conversations, product messaging, and SEO research into narratives that shape how organizations understand SIEM, threat detection, and cloud security, with a strong focus on clarity and real-world relevance.

Want ITSM best practice and advice delivered directly to your inbox? Why not sign up for our newsletter? This way you won't miss any of the latest ITSM tips and tricks.

nl subscribe strip imgage

More Topics to Explore

Leave a Reply

Your email address will not be published. Required fields are marked *