SLA Escalation Playbooks

SkillProductivity

A cross-PSA escalation framework for SLA pressure: how each PSA family (Autotask, HaloPSA, ConnectWise Manage, Syncro, Kaseya BMS) models SLA/priority state and where breach risk lives in each, a normalized breach-risk state model (healthy, at risk, breached-response, breached-resolution) with the default escalation action per state, how notification audience shifts by contract tier, and the evidence to gather before paging anyone.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the SLA Escalation Playbooks skill

What this skill tells your AI

The instructions your AI receives, as published by wyre-ai/msp-claude-plugins in msp-claude-plugins/ops-pack/skills/sla-escalation-playbooks/SKILL.md and read by ahel’s review.

Overview

SLA breaches are the single most reliable predictor of a client escalation call. This skill is the judgment layer on top of whatever SLA/priority data your connected PSA returns: it tells you when a ticket's breach risk crosses a threshold that warrants action, who should be notified at that threshold, what evidence to attach to the escalation, and how the response changes for a Platinum client versus a break-fix client on the same board.

This skill does not re-teach any single PSA's ticket API. It assumes you already know how to fetch a ticket and its SLA fields via the connected vendor's tools (or via conduit__search_tools if you don't yet know which tools are available) — what it adds is the cross-vendor decision logic for what to do with that data once you have it.

Anti-triggers

  • A platform's SLA policy configuration — policy fields, business-hours calendars, and how a vendor computes its own due dates are that connector's surface; use freshdesk-sla-business-hours, autotask-tickets, or halopsa-tickets.
  • Engineering error budgets — SLO burn rate against an uptime or error-rate target is a different measurement from a contractual response time; use error-budget-tracking in devops-pack.

Key Concepts

How each PSA family models SLA state

Every PSA expresses "how close is this ticket to breaching" differently. Resolve the concrete field/tool names for the connected instance before relying on any of this — these are the shapes to expect, not literal API contracts:

PSA familySLA/priority modelWhere breach risk lives
AutotaskNumeric ticket priority (1–4) plus a Service Level Agreement linked to the client's contract, driving separate first-response and resolution due-date fields on the ticketautotask__get_ticket_details / autotask__search_tickets — look for resolution plan / due-date fields; autotask__list_ticket_priorities resolves the priority label
HaloPSASLA profile assigned per ticket (often derived from client + priority), with explicit response and resolution target timestamps and a breach flaghalopsa__tickets_get returns deadlinedate / SLA hold state; halopsa__tickets_list can be filtered/sorted by SLA proximity
ConnectWise Manage/PSASLA record tied to board + priority, combined with Impact/Urgency fields; boards often carry their own escalation status flagTicket record's SLA/status fields — confirm the board's escalation flag naming, it is board-configurable
SyncroLighter-weight: ticket "Due Date" plus priority, no separate formal SLA engine in most instancesTicket due-date field and priority; treat due-date proximity as the SLA proxy
Kaseya BMSService Desk SLA tied to the client's Service Level Agreement, with response/resolution timers per ticketTicket SLA timer fields exposed by the connected Kaseya BMS tools

If the org has more than one PSA connected (rare, but happens during a PSA migration), scope explicitly to one board/instance per run and say which one you used.

The common escalation decision framework

Regardless of which PSA is behind the numbers, normalize every ticket to a single breach-risk state before deciding what to do:

  1. Healthy — comfortably inside both response and resolution targets.
  2. At risk — inside target but less than ~25% of the allotted window remains (tune this threshold to the org's own norms if documented; state the threshold you used).
  3. Breached — response — first-response target missed; no technician has substantively engaged yet.
  4. Breached — resolution — resolution target missed; ticket has had engagement but is not resolved.

Escalation action scales with state:

StateDefault action
HealthyNo action
At riskInternal nudge to the assigned technician (or to the dispatcher if unassigned)
Breached — responseEscalate to team lead/service manager; internal note logged on the ticket
Breached — resolutionEscalate to service manager and, per contract tier below, to the client

Evidence to gather before escalating

Never escalate on the SLA timer alone — attach the context a manager or client contact will actually need:

  • Ticket age and full status-transition history (when did it last move, and to what)
  • Assigned technician (or confirmation it's unassigned) and their current open-ticket count, if the PSA exposes technician workload
  • Last client-visible communication timestamp and its content
  • Whether the ticket is waiting on the client (see board-hygiene skill) — a breach caused by a non-responsive client is a different conversation than one caused by an idle queue
  • Contract/SLA tier for the client (see below)

Escalation by contract tier

Contract tier changes who gets notified and whether the client is proactively contacted, not whether the breach itself matters:

Tier (typical naming)On "at risk"On "breached"
Premium / Platinum / fully-managedTeam lead notified at "at risk"; account manager loop-in on breachProactive client contact before the client notices, with a remediation ETA
Standard / ManagedTeam lead notified on breach onlyClient contact only if they ask, but the internal note documents the breach for QBR reporting
Bronze / Block-hours / break-fixInternal reassignment onlyClient contact per the ticket's own status, no proactive SLA-specific comms (these contracts often don't carry a formal SLA at all)

Confirm actual tier-to-contract mapping against the PSA's contract/agreement records (or client-360-briefer-style client context if the org also has the wyre-gateway plugin's agents available) rather than assuming — tier names vary a lot by MSP.

If no PSA is connected

This skill cannot compute breach state without a source of ticket/SLA truth. If conduit__search_tools (or a direct attempt at a PSA tool call) shows no PSA connector present, say so explicitly and stop rather than guessing: "No PSA is connected, so SLA state can't be verified. Here is the general escalation framework above — apply it manually once ticket data is available." Do not fabricate SLA figures or invent a ticket's breach state.

Best Practices

  • Log the escalation itself as a ticket note/action where the PSA supports it, so the audit trail lives with the ticket, not just in chat.

Related Skills

  • Dispatch Prioritization — scoring and assigning the unassigned queue, of which SLA proximity is one factor
  • Board Hygiene — stale and stuck-ticket detection, including distinguishing "waiting on client" from a real SLA risk

Signals

GitHub stars
45
Forks
24
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
sla-escalation-playbooks
Source
github.com/wyre-ai/msp-claude-plugins
SLA Escalation Playbooks: Skill · ahel