A. Levin — Case File
Subject: Optimizing core operational process with AI Filed by: Alan Levin Practice: Service & incident management, 25 yrs

Service operations · incident & problem management · AI

Core operational processes should be optimized with AI — not just documented and left alone.

I'm Alan Levin. I've spent twenty-five years running service, incident and problem management for cloud platforms at Oracle, Microsoft, Ericsson, eBay/PayPal and WebEx. What I'm building now is the next step: encoding that operational judgment as tested AI skills, wired into the applications teams already run, so the process executes instead of sitting in a wiki.

Process plus applications plus AI — not one instead of another.

The process already exists

Most organizations have a severity matrix, an escalation path, and a postmortem template. The problem isn't that the process is missing — it's that under pressure people skip it, apply it inconsistently, or can't find it.

The applications already exist

Prometheus, Grafana, Kibana, PagerDuty, ServiceNow, Jira. The data needed to run the process well is already being collected. It just sits in six tools that don't talk to each other in the moment it matters.

AI is the connective layer

A model that reads live telemetry, applies the organization's own rules consistently, drafts the artifacts, and writes back into the system of record — while a human keeps the decisions that should stay human.

Built and in progress.

Each ITIL practice area gets the same treatment: real operational rules, encoded as skills, tested against realistic scenarios, and mapped to the tools that already run the workflow.

  • Built · 7 skills

    Service desk

    The front door: intake and classification, first-touch resolution, request fulfillment, escalation and handoff, user communication, closure and satisfaction, and performance review — with the key process indicators captured at each step.

    See the practice →
  • Built · 7 skills

    Incident & problem management

    Triage and severity classification, live diagnosis, multi-audience status updates, resolution and RFO, problem records, root cause analysis with CAPA, and known error records. Includes a runnable demo and the full integration architecture.

    See the practice →
  • Built · 7 skills

    Change enablement

    Intake and classification, risk and impact assessment, CAB authorization, scheduling and conflict detection, implementation and rollback planning, emergency change, and post-implementation review — plus how change connects to every other practice.

    See the practice →
  • Planned

    Service level management

    SLA breach notification, service review reporting, OLA alignment.

  • Planned

    Knowledge management

    Turning resolved incidents and problems into findable knowledge articles.

Twenty-five years running this for real.

35%reduction in time to mitigate across Oracle's SaaS portfolio
50%faster customer notification of impacting events
99.99%availability sustained through 10x demand growth at Ericsson

VP roles at Oracle across SaaS Production Engineering and the Oracle Applications Lab, where I established the Major Incident and Problem Management framework and serve as Incident Commander for critical issues. Before that, service engineering leadership at Microsoft, Ericsson, eBay/PayPal and WebEx.

Hiring for TPM or incident management roles

I'd like to talk about how this kind of work applies to your team.

ahlevin@hotmail.com
linkedin.com/in/alanlevin

Want this running for your team

Interested in these skills — or something like them — supporting how your team actually operates. Let's talk about what that would take.

ahlevin@hotmail.com
linkedin.com/in/alanlevin