When the System Works and the Service Still Fails

Illustrated desk scene with a stained official document representing a service failure caused by an erroneous death report

Summary

Mark Patnaude’s wife sadly learned she was “dead.” An erroneous death report had entered the Social Security Administration’s system at the state level and propagated for five weeks before anyone knew to challenge it – during which time the incorrect status had already affected multiple downstream services. Every interface ran. Every transaction completed. Nothing on the operational dashboard flagged anything wrong. Mark’s argument, drawn from that experience and grounded in ITIL (Version 5) thinking on outputs, outcomes, and the “focus on value” guiding principle, is that technical performance metrics and service success are not the same thing, and the gap between them is exactly where the most serious failures hide.

Systems can work as designed, and the customer can still experience the wrong outcome. The service failure experience below reminded me that measuring technology performance isn’t always the same as measuring service success – the system can work, but the service still fails.

A Service Failure Example

A few days ago, my wife learned that she was dead.

Not metaphorically.

An erroneous death report had reached the Social Security Administration identifying my very-much-alive wife as deceased. There was just one obvious issue: she was standing next to me when we learned about it.

We later determined that the erroneous report originated at the state level on July 3 and was subsequently entered into the federal system. We did not discover what had happened until August 7 – five weeks later.

During this time, the incorrect status had already affected multiple systems and services that relied on the information. From our perspective as customers, those consequences were largely occurring out of sight. We didn’t know there was an issue to challenge because we didn’t know the issue existed.

Once the error was confirmed, we were told that correcting the record through what the office referred to as the “reincarnation” process would take at least another two weeks.

I don’t know every technical interface the record passed through, and I won’t speculate about where a validation control should have caught the error. What I do know is that incorrect information crossed organizational boundaries, was treated as authoritative, and produced real consequences long before the customer knew anything was wrong.

After spending much of my career leading technology and operations, this was the part I couldn’t stop thinking about.

The technology may have been working exactly as designed. It was still a service failure.

When Successful Processing Produces the Wrong Outcome

We often measure technology from inside the technology organization. Was the application available? Did the interface run? Was the transaction processed? Did the data synchronize? Was the incident resolved within the service level agreement (SLA) target?

Those are important questions. I’ve spent years working with these kinds of operational measures. But this experience reinforced another question that is much easier to overlook: did the customer receive the intended outcome?

Consider what happens when incorrect information enters an interconnected environment, and downstream systems act on it.

The interfaces may work. The transactions may complete. The applications may remain available. No infrastructure alarm necessarily goes off.

Bad information moving successfully between systems isn’t a successful service. It’s a service failure despite operating efficiently.

An integration that stops working is visible. An integration that successfully moves incorrect information can be much harder to detect.

How Long Can a Service Continue Producing the Wrong Outcome? 

Our experience put a number on the issue: approximately five weeks passed between the erroneous report and our discovery of it. This raises another service management question: how long can a service continue producing the wrong outcome before either the provider or the customer knows something is wrong?

Traditional monitoring is very good at detecting unavailable systems, failed interfaces, processing errors, and exceeded thresholds. It is much harder to alert on a transaction that completes successfully using information everyone assumes is correct.

The customer, however, eventually experiences the result.

My wife wasn’t concerned with which database contained the information, which interface transmitted it, or which organization owned each part of the process. She experienced one outcome: systems were treating her as dead when she wasn’t.

This is what makes the ITIL principle of think and work holistically so important. Customers don’t experience our organizational charts, technology stacks, vendor boundaries, or support structures. They experience the result produced by all of them together.

When Everything on the Dashboard Is Green

ITIL (Version 5) gives us useful language for examining this distinction: utility, warranty, outputs, and outcomes.

Utility and warranty help us examine whether a service is fit for purpose and fit for use. Outputs are what activities produce. Outcomes are the results stakeholders experience or seek to achieve.

The distinctions matter because they keep us from treating successful technical activity as synonymous with successful service delivery.

A process can produce its expected output. An interface can transfer information. A transaction can complete on time. Individual technology components can meet their performance and availability requirements.

None of that, by itself, guarantees the intended outcome, i.e. it’s a service failure.

Our experience was an extreme example. Information was accepted and acted upon while the stakeholder outcome was completely wrong.

An output tells us something happened. An outcome tells us whether what happened mattered in the way we intended.

A Service Failure Despite Dashboards Being Green

Imagine an operational dashboard showing:

  • Application availability above target 
  • Overnight processing completed 
  • Integration jobs successful 
  • No infrastructure alerts 
  • Transactions processed within SLA.

Everything is green.

Meanwhile, a customer is standing in an office proving that she is alive.

This isn’t an argument against operational metrics. Organizations need them. The issue comes when measures of technical performance become substitutes for measures of service outcomes.

An SLA can tell us whether an agreed target was achieved. A monitoring platform can tell us whether technology is operating within expected parameters. Neither automatically tells us whether the customer received value.

That’s where focus on value becomes more than an ITIL guiding principle printed on a slide.

Sometimes the most important question behind a green dashboard is simply: What aren’t we measuring?

Detection Is Part of the Service

The timeline also changed how I thought about the incident.

Most technology organizations pay close attention to how quickly they restore service once something goes wrong. Mean time to restore (MTTR), incident aging, SLA attainment, and escalation times are familiar measurements.

But restoration starts only after someone knows there is an issue.

In our case, the erroneous status existed for five weeks before we knew enough to challenge it. Once we did, we were told the correction process could take at least another two weeks.

This separates two very different questions: How quickly can we correct a service failure once it is known? And how quickly do we know that the customer is experiencing a service failure in the first place?

A service can appear healthy internally while a customer is experiencing an increasingly serious issue externally.

That makes detection, notification, transparency, and the ability to challenge incorrect information part of the service experience – not merely administrative details surrounding it.

Fixing the Record Isn’t the End of the Issue

For the affected person, the immediate priority is obvious: correct the information and restore any services disrupted.

From a service management perspective, this is only part of the issue.

The experience raises broader questions. How did incorrect information become trusted? What validation occurred before consequential actions were taken? Which other services relied on the same information? How easily can someone challenge a record that everyone else’s systems believe is correct?

And when a service failure crosses several systems or organizations, who owns the customer’s overall outcome?

Correcting one record helps one person. Understanding what allowed incorrect information to be accepted and acted upon can improve the service for everyone who depends on it.

It also illustrates why data quality belongs in service management conversations.

Modern services increasingly depend on information moving between applications, providers, automated workflows, and external systems. The reliability of that information is part of the service’s reliability.

As organizations automate more decisions, this becomes even more important.

Artificial intelligence (AI) makes the issue especially visible, but the issue existed long before generative AI (GenAI) arrived. Giving systems greater ability to act on information increases the importance of understanding where that information came from, whether it can be trusted, and what happens when it is wrong.

Automation can speed up a good process. It can also make a bad outcome arrive faster.

Service Failure: Design for the Exception

Most services are understandably designed around what normally happens. Information enters a system, passes validation, triggers a workflow, and produces an expected result.

The more interesting service management question is what happens when reality doesn’t match the system’s expectation.

My experience left me with a question I’d encourage service leaders to ask about their own environments:

Where could a technically successful transaction create a completely wrong customer outcome?

That question leads quickly to others:

  • Which data elements can trigger significant downstream actions? 
  • Which systems treat information from another system as authoritative? 
  • What validation occurs before consequential actions are taken? 
  • How quickly would we know if trusted information were wrong? 
  • Can customers easily challenge information they believe is incorrect? 
  • Can employees intervene when evidence clearly contradicts the system? 
  • Can information be traced back to its source? 
  • When an error crosses multiple services, who owns restoring the customer’s overall outcome? 
  • Are we measuring successful transactions or successful outcomes? 

These aren’t purely technical questions. They cross service design, governance, risk, information management, suppliers, people, and customer experience.

They also expose something that becomes increasingly important as we automate more of our services: sometimes human judgment isn’t an inefficiency in the process. It’s an essential control.

The Need for Human Judgment

Eventually, people had to look at the evidence and the person standing in front of them, and recognize that the information in the system was wrong. No amount of confidence in the system could change the observable reality.

The objective shouldn’t be to preserve manual intervention everywhere. It should be to understand where automation creates value, where human judgment creates value, and where one needs to provide a check on the other.

The Question Behind the Dashboard

Technology leaders should continue measuring uptime, incidents, SLAs, transaction performance, and service availability. I certainly will.

But this experience reinforced two questions I think belong alongside all of them:

Did the service produce the outcome it was supposed to produce?

And if it didn’t, how quickly would we know?

Systems can be available, interfaces can run, transactions can complete, and SLAs can remain green.

And the customer can still be standing in an office proving that she is alive.

When this happens, the technology may have worked.

The service didn’t.

Service Failure FAQs

What is this article about? 

It’s a first-person account of Mark Patnaude’s wife being erroneously recorded as deceased in the Social Security Administration’s system, and the five weeks it took to discover the error. Mark uses the experience to explore a core service management idea: technical systems can function perfectly while still failing to deliver the outcome the customer actually needed.

What really happened to Mark Patnaude’s wife? 

An incorrect death report entered at the state level on July 3 was passed into the federal system. The error went undetected for about five weeks, until August 7, and once discovered, the correction process (referred to by the office as “reincarnation”) was expected to take at least two more weeks.

What is the difference between “outputs” and “outcomes” in this article? 

Outputs are what a process or system produces – a completed transaction, a successful interface run, a synchronized database. Outcomes are the actual results experienced by the stakeholder. The article’s central argument is that a system can generate all the right outputs while still producing the wrong outcome for the customer.

What is ITIL, and why does Mark Patnaude reference it?

ITIL (formerly known as the IT Infrastructure Library) is a widely used framework for service management. Mark Patnaude draws on ITIL (Version 5) concepts – including utility, warranty, outputs, and outcomes, as well as the guiding principles “think and work holistically” and “focus on value” – to explain why technical performance metrics alone don’t guarantee good service.

Who is this article written for? 

Primarily IT and service management leaders, ITIL practitioners, and technology operations professionals – though the underlying lesson about detection, data trust, and outcome-focused metrics applies to any leader responsible for customer-facing systems.

What’s the main takeaway or call to action? 

Mark Patnaude encourages service leaders to ask two questions: “Did the service produce the outcome it was supposed to produce?” and “If it didn’t, how quickly would we know?” He argues that detection and the ability to challenge incorrect information should be treated as part of the service itself, not just an administrative afterthought – and that human judgment remains an essential check on automated systems.

Does this article relate to AI? 

Yes, briefly. Mark Patnaude notes that AI and automation make the underlying issue more visible and more urgent, since automated systems can act on bad information faster – but he’s clear the issue existed long before generative AI.

Is this a technical or a leadership-focused article? 

It’s leadership-focused. While it references specific ITIL terminology, the article is written as a reflective, narrative piece for a general professional audience rather than a technical how-to.

Further Reading

Mark Patnaude
Mark Patnaude
Technology Leader

Mark Patnaude is a technology leader with more than two decades of experience in IT operations and service delivery. His focus is on making sure technology delivers real value to the business and the people it serves. He is ITIL (Version 5) Foundation certified.

Want ITSM best practice and advice delivered directly to your inbox? Why not sign up for our newsletter? This way you won't miss any of the latest ITSM tips and tricks.

nl subscribe strip imgage

More Topics to Explore

Leave a Reply

Your email address will not be published. Required fields are marked *