ITSM Knowledge Management

ITSM Knowledge Management: How Living Knowledge Drives Resilient ITSM and True Shift-Left

Spread the love

When I joined the IT force back in 2008, managing incidents was part of my core job as an Application Support executive.

I joined a team with professionals with a lot of experience in the field, and luckily, the top Indian IT company I joined was big on process. In the absence of this strong base early in my career, my ITSM journey would have been very different.

One of the extremely essential components in a strong ITSM system is Knowledge Management. Every single action we take as a response to a user-raised request or a service outage should be backed by either an existing knowledge base or have a post-incident review process, where we create documents for future reference. This is what strengthened my incident management game and helped me be ready for any outage.

When I moved to other companies and projects after my first experience, I realised that what I considered to be inherent to the system, and second nature for an ITSM professional, was not consistent across projects and was largely driven by the people involved.

I vividly remember the chaos I experienced 2 weeks into a new project, where the only thing I could gather was that the issue was nobody’s fault, since every team tried to push the fault on another.

I noticed that more often than not, such knowledge base articles rot in archives, built once but never updated. In such organisations, I noticed that when an outage hits, the analysts are scrambling through obsolete runbooks, guessing workarounds or escalating to the L2 or L3 support staff.

This results in high Mean Time To Resolution (MTTR), recurring incidents and burnout.

Static Documentation – ITSM Flaw

Static documentation, i.e., the documentation that gets created once and forgotten as just another task, is as bad, if not worse, than technical debt. Similarly, any documentation which cannot be easily retrieved is equally useless.

Documentation should continuously evolve with every software version change and feature update.

We cannot have the evolving project design and bug-fix knowledge stay in the heads of people who implement the change.

One of the root causes that people hate documentation is that it feels repetitive and something that gets outdated quickly. In the absence of a predefined system, it is also often cumbersome to maintain or enhance existing knowledge.

Overall, managing documentation seems like a boring and thankless job. But I personally think a boring system where everything just works is better than the excitement of a Sev 1 outage where all hell breaks loose.

Preventative measures in place via robust documentation ensure we do not need heroes – something I referred to in this post.

Knowledge Management in Agile World

Some may believe that in a fast-evolving Agile World, the knowledge management process would only slow everything down.

Are we moving faster if we keep solving the exact same outage every third sprint?

The fact is, in today’s fast-evolving IT landscape, documentation has become even more important than before. There is no point in being fast if you keep fixing the same issues every other sprint.

Luckily, we have very advanced AI tools available to do this for us without impacting productivity.

These tools keep an eye on code changes and functional changes and can provide required documentation in the blink of an eye.

There was a time when for a small team, making a resource available for documentation was not realistic, but advancement in AI has ensured that it is a non-issue now.

Knowledge – ITSM Immune System

An evolving knowledge base is like an evolving immune system. Strengthening it should be a top priority.

If we are aware of potential risks of a change and document it beforehand, it works the same as a flu shot in the flu season.

ITSM Knowledge Management is like vaccines

The way a flu shot can potentially improve recovery time, or avoid the flu altogether, a living knowledge base can help in ensuring the service desk staff know exactly what to do when an outage hits. An outdated knowledge document is as irrelevant as an outdated flu shot.

For outages that were not factored in for whatever reason, the frontline engineers should ensure that the knowledge base gets updated with new knowledge, update verified workarounds and the resolution playbook, then, depending on tools available, embed it in the current incident management system.

Having this system in place helps not just the front desk professionals, but also the specialists further up the line.

Enabling True “shift left” In ITSM

Much like in the medical industry, even in IT, the top-of-the-line specialists are more expensive and fewer in number than frontline workers.

Much like how ensuring the frontline medical workers can make quick decisions based on basic triaging, it is essential to enable the service desk professionals to be able to implement workarounds and functional fixes in response to service requests, minor incidents, as well as server outages.

In a mature system, L1 agents should be able to resolve complex issues using validated methods handy to them, while L2 and L3 support can concentrate on Problem Management via root cause analysis without being interrupted by repetitive triage requests.

How much tribal knowledge leaves the room when your primary L3 specialist goes on leave?

Other than enabling the L1 agents to be more productive, robust documentation also ensures there is no dependency on one individual for a system to be maintainable.

Key Takeaway For ITSM Leaders

The Incident Management process is designed to restore service, but Knowledge Management ensures we do not have to solve the same issue twice from scratch.

If we want the system to be stable, MTTR to remain low, and avoid burnout, do not consider documentation as a chore but as the first line of defence against all risks.

As the IT systems expand, CI/CD pipelines evolve, and AI triaging becomes the next big thing, a culture that values documentation as much as solving issues not only provides necessary guardrails against a potential mishap, but also keeps your organisation running as a cohesive system.

Also, appreciate these nameless heroes who ensure you are prepared to tackle every issue the wild world of service management throws at you.

Hope this post gave you some insights and prompts you to audit and optimise your teams in the high-velocity world of IT Service Management.

Do subscribe below for more such ITSM strategy discussions.


Discover more from Jijo George

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *