Background
This came from running the process, not reading about it.
I've spent twenty-five years running service, incident and problem management for cloud platforms at Oracle, Microsoft, Ericsson, eBay/PayPal and WebEx — building the on-call programs, serving as Incident Commander on critical issues, and owning the numbers that come out the other side. At Oracle SaaS Production Engineering that meant cutting time-to-mitigate by more than 35% and halving the time to notify customers of impacting events.
So the severity matrix, escalation rules, and RFO/RCA distinctions in this suite aren't textbook ITIL. They're the actual judgment calls that come up when a Sev1 is live and someone needs an answer in the next fifteen minutes.
What's new is encoding that judgment as a tested, iterated set of Claude skills — the same draft → test → feedback → revise loop you'd use to build any other piece of software, applied to operational process instead of code.
Record of service
Alan Levin
Transformational leader with a track record of driving mission-critical initiatives supporting cloud services for Oracle, eBay, PayPal, Skype, Ericsson, Microsoft and WebEx. Hands-on builder and technology leader with deep expertise in service operations, incident and problem management, software delivery transformation, and technical program management.
-
VP, Technical Program Management
Oracle Mar 2023 – Present- Leads Oracle Applications Lab's (OAL) Program Management team driving automation and tooling programs that accelerate data center buildouts and enable Oracle's AI infrastructure growth.
- Established OAL's Major Incident and Problem Management framework, serving as Incident Commander for critical issues — proactive monitoring, structured response, Jira-based problem tracking, and AI-assisted root cause analysis driving corrective and preventive actions.
- Transitioned software delivery from SAFe to lightweight Agile, recovering over 10% of development resources in three months — annual resource recovery exceeding $10M alongside improved time to market.
- Built a Program and Portfolio Intelligence Platform with Codex and Oracle APEX to enforce governance standards and surface insights from execution through executive review.
- Led the Cerner operational integration, moving Quote-to-Cash from multiple third-party suppliers onto Oracle's Fusion stack in nine months, improving the QtC experience.
-
VP, SaaS Production Engineering
Oracle Mar 2018 – Mar 2023- Transformed a reactive Enterprise Operations organization into a proactive Cloud DevOps team running 24x7x365 incident response across Oracle's SaaS portfolio.
- Developed the incident management process and tooling framework for Oracle's SaaS portfolio — enabling self-healing and reducing time to mitigate by over 35% since inception.
- Drove the migration from siloed tools and teams for signal collection, alerting, correlation and response to an integrated full-cycle capability — and cut time to notify customers of impacting events by 50%.
- Restructured and modernized a 300-person organization while lowering costs by more than 50%, avoiding roughly $36M over three years.
- Eliminated 50% of inbound work request volume while improving first-touch resolution through automation and self-service.
-
VP, Service Delivery and SRE
Ericsson / MediaFirst Jan 2014 – Mar 2018- Established the DevOps model and built a Global Service Center for B2B2C media solutions, launching MediaFirst as a V1 SaaS TV platform.
- Hired and led a global team of 40+ service engineers, SREs, CI/CD engineers, program managers and architects supporting 24x7x365 TV solutions.
- Achieved >99.99% availability through 10x demand growth via proactive surveillance, automation, HA cloud architecture and SRE discipline.
- Implemented an Agile release model — CI/CD, feature toggles, blue/green releases — that cut time to market by over 50% at carrier-grade availability.
-
Director, Service Engineering and Hardware Delivery Solutions
Microsoft Mar 2006 – Jan 2014- Led budget, strategy and operations for a 24x7x365 Service Operations Center providing monitoring plus incident, change, release and problem management for Microsoft cloud services.
- Managed integration and service launches for acquisitions including FAST Search, Skype and Yammer; led disaster recovery remediation for a business generating more than $1M per month.
- Reduced infrastructure delivery cycle time by more than 50% and lifted customer satisfaction from 66% to 95%.
- Scaled cloud deployment services through 5x year-over-year demand growth across Azure, Xbox Live, Messenger and Windows Live ID.
-
Manager, Service Operations
eBay / PayPal Jul 2004 – Mar 2006- Oversaw change, release and problem management while developing processes and tools for eBay.com and PayPal.com site operations.
- Program managed PayPal's Availability Steering Committee, driving availability above 99.95%.
- Built a capacity planning and infrastructure delivery process that eliminated capacity-driven outages by identifying risks months ahead of service impact.
- Created disaster recovery plans and supported third-party and internal audits including SOX, FSA and ACH.
-
Sr. Manager, Service Operations
WebEx Jun 2000 – Jul 2004- Managed data center operations supporting 24x7x365 monitoring and deployment of a cloud-based SaaS solution.
- Raised service availability from 99.95% to 99.99% by working across engineering, product management and customer care.
- Built service provisioning and onboarding capabilities to scale through 80x customer growth.
- Pioneered offshore operations and automation, cutting cycle times by more than 50% and avoiding over $1M annually.
-
Strategic Supply Planning
Intel Corporation Jan 1996 – Jun 2000- Developed product build strategies and what-if models for Intel Architecture microprocessors, supporting build plans exceeding 100M units annually.
- Automated build-plan analysis, halving monthly planning effort and saving 40 hours per week at each of five Intel sites.
Carnegie Mellon University — B.S., Industrial Management. Captain, Varsity Soccer.
Hiring for TPM or incident management roles
I'd like to talk about how this kind of work applies to your team.
ahlevin@hotmail.comlinkedin.com/in/alanlevin
Want this running for your team
Interested in these skills — or something like them — supporting how your team actually operates. Let's talk about what that would take.
ahlevin@hotmail.comlinkedin.com/in/alanlevin