Service operations · incident & problem management · AI
Core operational processes should be optimized with AI — not just documented and left alone.
I'm Alan Levin. I've spent twenty-five years running service, incident and problem management for cloud platforms at Oracle, Microsoft, Ericsson, eBay/PayPal and WebEx. What I'm building now is the next step: encoding that operational judgment as tested AI skills, wired into the applications teams already run, so the process executes instead of sitting in a wiki.
The idea
Process plus applications plus AI — not one instead of another.
The process already exists
Most organizations have a severity matrix, an escalation path, and a postmortem template. The problem isn't that the process is missing — it's that under pressure people skip it, apply it inconsistently, or can't find it.
The applications already exist
Prometheus, Grafana, Kibana, PagerDuty, ServiceNow, Jira. The data needed to run the process well is already being collected. It just sits in six tools that don't talk to each other in the moment it matters.
AI is the connective layer
A model that reads live telemetry, applies the organization's own rules consistently, drafts the artifacts, and writes back into the system of record — while a human keeps the decisions that should stay human.
Practice areas
Built and in progress.
Each ITIL practice area gets the same treatment: real operational rules, encoded as skills, tested against realistic scenarios, and mapped to the tools that already run the workflow.
-
Built · 7 skills
Service desk
The front door: intake and classification, first-touch resolution, request fulfillment, escalation and handoff, user communication, closure and satisfaction, and performance review — with the key process indicators captured at each step.
See the practice → -
Built · 7 skills
Incident & problem management
Triage and severity classification, live diagnosis, multi-audience status updates, resolution and RFO, problem records, root cause analysis with CAPA, and known error records. Includes a runnable demo and the full integration architecture.
See the practice → -
Built · 7 skills
Change enablement
Intake and classification, risk and impact assessment, CAB authorization, scheduling and conflict detection, implementation and rollback planning, emergency change, and post-implementation review — plus how change connects to every other practice.
See the practice → -
Planned
Service level management
SLA breach notification, service review reporting, OLA alignment.
-
Planned
Knowledge management
Turning resolved incidents and problems into findable knowledge articles.
Track record
Twenty-five years running this for real.
VP roles at Oracle across SaaS Production Engineering and the Oracle Applications Lab, where I established the Major Incident and Problem Management framework and serve as Incident Commander for critical issues. Before that, service engineering leadership at Microsoft, Ericsson, eBay/PayPal and WebEx.
Hiring for TPM or incident management roles
I'd like to talk about how this kind of work applies to your team.
ahlevin@hotmail.comlinkedin.com/in/alanlevin
Want this running for your team
Interested in these skills — or something like them — supporting how your team actually operates. Let's talk about what that would take.
ahlevin@hotmail.comlinkedin.com/in/alanlevin