CAST AI Failure Triage
SkillAI & modelsDiagnose CAST AI connection, agent, node autoscaling, and workload autoscaling failures without making speculative changes. Use when a cluster is disconnected, recommendations are absent, pods are not optimized, or capacity does not scale as expected. Trigger with: "debug CAST AI", "CAST AI is not scaling", "why is CAST AI disconnected".
Use CAST AI Failure Triage in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add CAST AI Failure Triage and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the CAST AI Failure Triage skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by jeremylongshore/tons-of-skills-marketplace in skills/.curated/castai-common-errors/SKILL.md and read by ahel’s review.
Overview
Separate observation, connectivity, policy, capacity, and disruption failures before proposing a change. Preserve the failing state, use current component topology, and stop when the evidence requires cloud-provider or CAST AI support access.
Prerequisites
- The exact kube context, cluster, region, time window, and observed symptom
- Read-only access to the
castai-agentnamespace - The declared installation owner: castctl, Terraform, GitOps, or console
Instructions
Step 1: Freeze the symptom
Record expected versus actual behavior, timestamps, workload identity, pending-pod reason, and recent configuration changes. Use Read and Grep on runbooks and IaC to determine whether Cost Monitoring, Node Autoscaling, or Workload Autoscaling is actually enabled.
Step 2: Check installation health
Use Bash(castctl:) for version or non-mutating status commands supported by the installed client. Use Bash(helm:) to inspect releases and values, then Bash(kubectl:*) to inspect workloads, readiness, events, and bounded logs in castai-agent. Do not restart components before collecting evidence.
Step 3: Classify the failure plane
| Plane | Evidence | Likely boundary |
|---|---|---|
| Connection | Agent readiness, outbound failures, console disconnect | Identity, network, or cloud permissions |
| Node scaling | Pending pods, policy bounds, node-template fit | Unsatisfied constraints or maximum CPU boundary |
| Workload scaling | Missing recommendations, policy assignment, metrics | Metrics server, confidence, policy, or unsupported workload |
| Disruption | Eviction denial, PDB events, deferred changes | PDB or selected apply mode |
| Reporting | Missing cost or savings window | Ingestion, baseline, adoption, or pricing configuration |
Step 4: Test one hypothesis
Choose the smallest reversible check. Confirm regional endpoint alignment, effective scaling-policy assignment, metrics availability, supported workload type, node-template constraints, and cloud quota. Treat the deprecated cluster minimum CPU setting as migration debt, not a current control to add.
Step 5: Decide the owner and remedy
Map the evidence to the owning layer. Change repository-managed values only through their source of truth; do not mix console edits into Terraform or GitOps ownership. Escalate with a redacted bundle when the failure is inside the hosted control plane or an undocumented provider response.
Tool Discipline
Use Read and Grep for configuration and runbook evidence. Use Bash(kubectl:), Bash(helm:), and Bash(castctl:*) only for bounded inspection commands. Do not apply, upgrade, restart, connect, disconnect, or expose Secret objects during diagnosis.
Output
- A timestamped symptom and environment summary
- Evidence grouped by failure plane
- One supported root-cause hypothesis with confidence
- A reversible remedy, rollback condition, and escalation owner
Examples
Recommendations are absent because metrics-server is missing, so the remedy belongs to cluster observability. A node remains pending because every approved node template conflicts with its constraints; increasing a global limit without reviewing the workload is not the remedy.
Error Handling
| Failure | Response |
|---|---|
| Kube context is ambiguous | Stop before any cluster command and resolve it |
| Logs include credentials or inventory | Redact locally and do not attach raw output |
| A PDB blocks Immediate mode | Preserve the PDB and evaluate Deferred mode with the workload owner |
| Evidence points to cloud quota | Escalate to the cloud owner with the exact denied dimension |
Resources
Signals
- GitHub stars
- 3k
- Forks
- 415
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
castai-common-errors- Source
- github.com/jeremylongshore/tons-of-skills-marketplace
github.com/jeremylongshore/tons-of-skills-marketplace
Related picks
Skill · microsoft
The pick for Kubernetesdt-obs-kubernetes
Skill · dynatrace
The pick for Kuberneteshelm-chart-scaffolding
Skill · davila7
The pick for Helminfra-containers-kubernetes
Skill · agents-inc
The pick for Kubernetesvigilante-issue-implementation-on-terraform
Skill · aliengiraffe
The pick for Terraforminfra-iac-terraform
Skill · agents-inc
The pick for Terraform