Is Hero Culture In IT A Ticking Time Bomb?
Table of Contents
What is Hero Culture In IT?
Hero Culture in IT has been a thing I have experienced in every organisation I have been part of.It may be more in some organisations than others, but it is a common case across industries.
This was early in my career, and I think it was during my annual performance review meeting after the ratings were done. I was obviously not happy with how I was rated and decided to talk to my manager about it.

I was told, “You were up there with another person to get the top rating, but she was chosen because she recently helped out during a critical outage and impressed everyone when the deployment failed and led to a P1 outage.”
My next question was,
“Was she in operations too?”
The response:
“She was the developer.”
“So you rewarded someone who deployed broken code to production, instead of someone who has been ensuring no broken code hits production and making sure all preventative measures are in place?”
I was told, “Unfortunately, this is how it is.”
This was my first experience with how the heroes of ITSM are usually not the ones who prevent disaster, but the ones who swoop in at the last minute, in full public display, to fix the problem (even if it was they who caused it).
Why Are Heroes Hailed in IT?
It is easy to guess why such heroics are rewarded more than the efforts of those who follow SOPs and prevent disasters. A disaster brings visibility.
A web server that never fails has a team of ITSM-literate hardware and software engineers who keep it running. Much like in any other industry, these silent heroes are underappreciated and ignored. It is only when the system fails that we notice them, often blaming them for the failure. For such teams, the last resort is for one of them to fix a bad patch and be hailed as a hero so they can showcase how talented they are.
Think back to the last Sev 1 call you were on—the one where the whole system was down. Did you also have the whole call waiting for that one person who would come in, fix the issue, and then be hailed as a hero?
I am not saying that these heroes should not be appreciated; of course they should. But at the same time, a good system is not one that merely gets fixed when it breaks, but rather one that does not experience outages at all.
This is why the role of preventative action should be appreciated and rewarded.
Heroics to Prevention in ITSM
The transition from a team that operates in BAU mode and hails heroes to a system that prioritises prevention can take time. This is especially true for teams that are set in their ways. A team that has operated like this for a decade may see these preventive actions as a threat to their routine. Preventing issues before they happen, when not appreciated, can make the team feel undervalued.
New ideas are often dismissed because they seem to challenge the status quo. Questioning the existing approach is somehow made to feel wrong, which does not bode well in the face of organisational politics.
What the teams involved need to realise is that this is not about questioning existing methods, but rather transitioning to a new model alongside changes in time and technology.
The waterfall model of development gave way to Agile, making it the new buzzword, but many organisations did a half-hearted job with this transition. The system thought in waterfall but executed in Agile. By that, I mean an Agile model does not mean fixing things only as they break; it means planning with a baseline to ensure a quick fix can be modular without breaking the entire system.
In the pre-GitHub era, parallel work was not easy: files were locked for editing, overlaps were common, and a constant connection was required to check things in or out. With GitHub, merging and reconciling code has become much easier. Yet, we still have organisations that are stuck in old ways because of the fear that a transition might break the system.
Technical Debt in ITSM
One of the major roadblocks in this transition is the unpredictability of change and how it can potentially break what is currently working. In IT, we are often told not to fix what is not broken, but people forget that these are simply systems that have not broken yet. Patches intended to prevent future issues are often bypassed with quick workarounds to avoid larger changes or scheduled maintenance. This only continues to pile on technical debt.
A major example of such debt leading to catastrophe was Knight Capital Group, an American financial services firm. Technical debt from obsolete, inactive code caused a loss of $440 million USD in just 45 minutes, driving the firm into a forced acquisition. You can read more about it here.
Failing to take preventative action can cause technical debt to explode into much larger issues. You may have a hero who can jump in and resolve the problem, but a one-hour IT outage can cost anywhere between $300k and $5 million depending on the industry—a loss that many mid-sized companies cannot afford.
How to Reward Prevention
Unlike reactive actions taken after an outage, which an organisation routinely tracks, the results of preventative action cannot always be directly measured. Something that cannot be measured is not easy to reward. There is no MTTR for a preventative action. In order to reward this approach, the metrics around ITSM need to change as well.
This is something I witnessed in organisations where I worked in the past. While managing an Application Support team in the 2010s, we tracked the efficiency of the team based not only on the incidents resolved, but also on the problem tickets closed, which led to a measurable reduction in recurring issues.
While problem management is still roughly 70% reactive, this approach can be expanded by being more proactive in CAB meetings and during the initial planning phase, where relevant performance metrics are documented and reviewed before implementing a new change.
Better inter-team coordination across every phase ensures that risks are properly accounted for. This also prevents finger-pointing post-incident, since constructive input during this phase earns the contributing team credit.
Tracking these inputs and automating metric analysis can further optimise the process, and the teams driving these optimisations can be rewarded accordingly.
People often fear that a stable system means the team maintaining it will become irrelevant. This is simply not true. On the contrary, a stable system ensures more bandwidth for improvements and enhancements, which helps both the system and the team grow. A system based on this principle runs like a well-oiled machine, giving the people involved more opportunity to think outside the box and introduce upgrades that pay dividends in efficiency and capital.
Conclusion
A stable system where nothing goes wrong is not an excuse to downsize a team; rather, it is an opportunity to reward those who contribute to that stability. These team members can be best utilised by giving them the freedom to implement their insights across the broader system. The bandwidth gained is not merely about cutting costs, but about increasing value—building a platform your competitors will envy, not just for its efficiency, but for its advancement, because stability and bandwidth deliver innovation when properly nurtured.
Discover more from Jijo George
Subscribe to get the latest posts sent to your email.
