PeopleCert Community

The Culture of “Firefighting” in ITSM

The Culture of “Firefighting” in ITSM
# Service Management
# Leadership Insights

Why operational maturity means preventing fires, not just putting them out.

August 27, 2026
Yurguen Penaranda Thomas
Yurguen Penaranda Thomas
The Culture of “Firefighting” in ITSM

The Culture of “Firefighting” in ITSM

Introduction: The Hero at Two in the Morning

It is 2:00 a.m. A critical service has gone down. The team is working under pressure to restore the service as quickly as possible. A specialist manages to restore it in record time and receives recognition from the entire organization.
Without taking away from the specialist’s achievement in restoring the service in the shortest possible time, what no one asks is why the incident occurred in the first place. Situations like these should raise concerns and lead organizations to question whether they are fostering a culture of firefighting.
A culture of firefighting is not characterized by the existence of major incidents. It is characterized by an organization that invests most of its efforts in reacting to problems that have already materialized, while dedicating few resources to understanding their causes and preventing their recurrence.

When Urgency Becomes the Norm

While an organization can implement actions to reduce the likelihood of a crisis occurring, there will always be some possibility that an event of this nature will materialize.
The problem is not facing an occasional crisis. The problem is when crises become the normal way of working.
When urgency becomes the norm, the organization begins to operate in “reactive mode.” Priorities change constantly and unexpectedly, and emergency meetings become routine. It becomes common for people to stop what they had planned in order to address unexpected urgent situations, generally caused by recurring incidents.
Postponing planned activities to address emergencies creates delays in agreed delivery timelines, and if those deliveries are client-facing, it can lead to dissatisfaction or even financial penalties. On the other hand, to avoid such delays, the organization may need to incur overtime costs so that employees can handle both emergencies and their planned responsibilities within established deadlines.
When these situations become normalized, they create a negative environment in which employees are constantly on alert, waiting for the next emergency that will force them to put everything aside and focus entirely on the crisis at hand.

The Operational Hero

Unconsciously, organizational leaders may be promoting this culture of firefighting. This can happen by recognizing employees primarily for resolving urgent situations or by becoming dependent on specific individuals to handle certain types of incidents due to their specialized knowledge. In some cases, this even reaches the point where these individuals unofficially maintain permanent availability so they can be contacted whenever an emergency arises within their area of expertise.
From a human perspective, heroes within organizations are admirable. From a service management perspective, however, their existence often reveals operational risks that have not yet been resolved.

The Hidden Cost of Living in Firefighting Mode

The presence of operational heroes is often perceived as a strength. However, when an organization systematically depends on these individuals to maintain service stability, costs begin to emerge that are not always visible at first glance.

People Burnout

This creates a sense of overload, as constantly handling emergency situations is far more physically and mentally demanding than carrying out everyday operational activities. This is even more evident when the resolution of these emergencies depends on a small number of subject matter experts.
This environment can negatively affect employees, leading to burnout, reduced performance, and, in the worst cases, increased turnover as people can no longer tolerate the environment and choose to seek opportunities elsewhere.

Reduced Capacity for Improvement

This firefighting culture has a reactive focus, where time is spent resolving incidents and restoring services and infrastructure through temporary solutions (workarounds). However, there is little or no time available to investigate the situation and identify and address the root cause, thereby preventing recurrence.

Unstable Priorities

Urgent incident response constantly displaces important tasks that had already been planned, leading to missed deadlines, possible penalties, or the need to allocate additional resources (higher costs) to avoid non-compliance.

Loss of Trust

Users may begin to lose confidence in the organization because they perceive an unstable operation caused by constant service and system disruptions, as well as missed commitments due to employees having to prioritize emergency response.
In the case of a company that provides technology services, this loss of trust may result in customers choosing to stop consuming its services, negatively affecting the organization’s revenue.

Why Do So Many Fires Occur?

Organizations rarely develop a reactive culture intentionally. In most cases, this condition is the result of multiple weaknesses that accumulate across processes, controls, management practices, and decision-making. Below are some of these causes.

Reactive Management

Efforts are focused solely on addressing symptoms through practices such as Incident Management or Emergency Change Enablement, which implement temporary solutions to restore services as quickly as possible. However, the organization does not dedicate resources to proactive practices such as Problem Management or Continual Improvement, which focus on identifying the root cause of incidents and implementing improvement actions that prevent them from recurring.

Lack of Prevention

Historical incident data is not analyzed to identify trends among incidents of a similar nature. This type of analysis would make it possible to correlate similar incidents and subsequently use the Problem Management practice to identify the root cause.
In addition, a weak or non-existent Monitoring and Event Management practice may prevent the identification of alerts that could indicate potential incidents, making it impossible to address the situation before it becomes a major incident.

Weak Change Evaluation

A deficient Change Enablement practice that does not perform adequate impact analysis may result in incidents caused by change implementations because insufficient actions were taken to reduce the risk of service disruption.

Weak and Poorly Automated Processes

Processes may exist, but at a low level of maturity. For example, the Incident Management practice may not have clearly defined escalation paths, leading to uncertainty regarding when and to which team an incident should be escalated if it cannot be resolved by the initial support levels. As a result, the duration of the service disruption continues to increase.
Additionally, relying on highly manual processes can hinder effective communication among the different teams involved in incident resolution.

Dependency on Individuals

A culture may exist where knowledge is concentrated among the most experienced individuals. Without a mature Knowledge Management practice, organizations become dependent on the involvement of these experts, which can increase resolution times if they cannot be contacted quickly.

Signs That an Organization Is Living in Firefighting Mode

The following are some indicators that an IT organization has a reactive culture focused on firefighting:
  • Recurring incidents are frequent.
  • Priorities change constantly.
  • Root cause analyses are rarely completed.
  • Corrective actions remain open for months.
  • The same problems reappear repeatedly.
  • Certain individuals are considered “indispensable.”
  • Most meetings are focused on urgent issues.

How to Measure a Reactive Culture

Through key indicators, it is possible to identify more objectively whether an organization has a reactive incident management culture. The following metrics can help detect this situation:
  • Percentage of recurring incidents: Helps identify incidents that occur repeatedly, making them strong candidates for analysis through the Problem Management practice in order to identify and address their root causes.
  • Number of emergency changes: Helps identify which incidents were resolved through the implementation of a change, allowing a deeper analysis of their context.
  • Average problem closure time: Helps assess how long problem investigations take to reach root cause identification, indicating whether the organization is giving sufficient importance to this type of analysis.
  • Number of overdue corrective actions: Helps determine whether the organization is actually dedicating time to implementing identified corrective actions or if they are simply documented out of routine without receiving the necessary attention.
  • Overtime hours associated with incidents: Helps identify how many additional working hours employees had to dedicate to handling emergencies while also completing planned activities in order to avoid missing established deadlines.

What ITIL Is Really Trying to Promote

ITIL promotes IT service management focused on value creation and the delivery of outcomes to customers. From a service reliability perspective, what customers primarily want is not for providers to respond quickly to incidents, but rather to reduce the number of incidents that can affect the performance of services and systems.
To achieve this, organizations must shift from a reactive philosophy of “putting out fires” to a proactive approach focused on “preventing incidents from occurring.” This can be achieved through the proper application of several ITIL management practices:
  • Monitoring and Event Management: Provides early warning of events and enables corrective actions before customers even notice an impact.
  • Measurement and Reporting: Enables the analysis of historical data related to recurring or high-impact incidents so that this information can later be evaluated through the Problem Management practice.
  • Problem Management: Aims to identify and address the root cause in order to prevent recurrence.
  • Incident Management: Clearly defining escalation matrices that establish when and to which team an incident must be escalated ensures incidents do not remain unnecessarily with teams that lack the technical capability or specialized knowledge to resolve them.
  • Change Enablement: Ensures that proper impact analysis is conducted, reducing the likelihood of incidents resulting from change implementation.
  • Knowledge Management: Helps reduce dependency on individual experts by providing operational guides that explain in detail how to resolve certain types of incidents, thereby reducing the duration of service disruptions.
  • Continual Improvement: Enables organizations to identify opportunities for process improvement, whether by simplifying processes, implementing stronger controls, improving process understanding among participants, or introducing automation.

Conclusion: Maturity Is Not Demonstrated During a Crisis

Organizations often admire those who put out fires. However, true operational maturity is not reflected in how quickly an organization responds to a crisis, but in its ability to prevent the same crisis from happening again. A mature organization does not need more heroes; it needs fewer fires.
To achieve this, a change in mindset is essential. Organizations must stop viewing analysis and investigation activities as a waste of time and money. Preventive practices must be given the same level of priority traditionally assigned to emergency response, fostering a culture focused on preventing incidents, sharing knowledge, and promoting innovation and continual improvement.
Sign in or Join the community
Where conversation, connection, and real-world practices come together.
PeopleCert Community
Create an account
Where conversation, connection, and real-world practices come together.
Comment (1)
Popular
avatar

Dive in

Related

Blog
Beyond KPIs: The Cobra Effect and the Watermelon Effect in ITSM
By Yurguen Penaranda Th... • May 26th, 2026 Views 76
Blog
The New ITIL Version 5: What It Means in the Real World of Digital Services
By Scott Everett • Feb 5th, 2026 Views 154
Blog
Culture as the True Success Factor in ITIL Adoption
By Yurguen Penaranda Th... • Jul 21st, 2026 Views 49
Blog
The Risk of Confusing Compliance with Value Creation in ITIL
By Yurguen Penaranda Th... • Aug 11th, 2026 Views 34
Blog
Beyond KPIs: The Cobra Effect and the Watermelon Effect in ITSM
By Yurguen Penaranda Th... • May 26th, 2026 Views 76
Blog
Culture as the True Success Factor in ITIL Adoption
By Yurguen Penaranda Th... • Jul 21st, 2026 Views 49
Blog
The Risk of Confusing Compliance with Value Creation in ITIL
By Yurguen Penaranda Th... • Aug 11th, 2026 Views 34
Blog
The New ITIL Version 5: What It Means in the Real World of Digital Services
By Scott Everett • Feb 5th, 2026 Views 154