Debug OpenShell Gateway Deployment
SkillCloud & infraHelps your agent diagnose why an OpenShell gateway deployment is failing, unreachable, or cannot start sandboxes.
Use Debug OpenShell Gateway Deployment in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Debug OpenShell Gateway Deployment and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Debug OpenShell Gateway Deployment skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Debug why an OpenShell gateway deployment is unhealthy, unreachable, or unable to create sandboxes. Use for gateway health failures, Docker/Podman runtime issues, Helm failures, Kubernetes scheduling, TLS or auth, gateway interceptors, supervisor middleware startup or runtime failures, external comp
What this skill tells your AI
The instructions your AI receives, as published by nvidia/openshell in skills/debug-openshell-cluster/SKILL.md and read by ahel’s review.
Diagnose a gateway and its selected compute platform. Do not assume OpenShell provisions Kubernetes or runs a k3s container. OpenShell targets a reachable gateway endpoint backed by Docker, Podman, Kubernetes, the experimental VM driver, or an operator-managed out-of-tree compute driver.
Use openshell first to identify the active endpoint. Then use the platform tools that match the gateway's compute driver: docker, podman, kubectl/helm, or VM driver logs.
Overview
The target deployment flow is:
- Operator starts or deploys the gateway with system packages, systemd, or Helm. The CLI does not start, stop, or destroy gateway services.
- Operator configures the compute driver.
- Operator provides the CLI and supervisor authentication material required by the deployment mode: edge or OIDC user auth, optional CLI mTLS, and gateway-minted sandbox JWTs.
- The CLI registers a reachable gateway endpoint with
openshell gateway add. - The gateway creates sandboxes through the selected compute driver.
The openshell-gateway composition crate explicitly installs its compiled
Docker, Podman, Kubernetes, and VM registrations at startup; openshell-server
does not link compute-driver crates. Custom gateway binaries may include a
subset of those registrations. With no configured driver, the gateway probes only
installed registrations in priority order (Kubernetes, Podman, then Docker);
VM has no probe and remains opt-in. Confirm the binary's registered drivers
when auto-detection reports that no suitable driver is available. If
configuration selects a driver that was not compiled in, the gateway treats
the name as an external driver and reports a missing socket_path unless an
endpoint is configured.
On Windows, custom binaries can include MXC independently. Registrations for Docker, Podman, Kubernetes, and VM are rejection stubs when included; they do not enable those runtimes on Windows.
See the compute driver reference for selective-build options and external-driver configuration.
For local evaluation only, TLS may be disabled and the gateway can be reached through http://127.0.0.1:<port>.
Prerequisites
- The
openshellCLI must be available for endpoint checks. - Know the active gateway name and endpoint, or be able to inspect local gateway metadata.
- Know the compute platform: Docker, Podman, Kubernetes, VM, or an out-of-tree driver.
- For Kubernetes:
kubectlmust target the cluster that hosts OpenShell and Helm version 3 or later must be available. - For Docker or Podman: the runtime socket must be reachable from the gateway host.
Use openshell --help and nested --help output as the authority for the installed CLI version. Use the published installation guide, compute-driver reference, gateway configuration reference, and Kubernetes setup guide as the authority for deployment and configuration behavior.
Workflow
Run diagnostics in order and stop once the root cause is clear.
Step 1: Check CLI Reachability
openshell gateway list --output json
openshell gateway info
openshell status
For a one-off endpoint check that bypasses stored gateway selection and metadata:
openshell --gateway-endpoint <url> status
Common findings:
No active gateway: register one withopenshell gateway add <endpoint>.- Connection refused: gateway process is not running, service exposure is wrong, or a port-forward/proxy is not active.
- TLS/certificate errors: the endpoint scheme or trust chain is wrong, a local mTLS bundle does not match the gateway CA, or TLS termination does not match the gateway listener.
- A Snap refresh restarts the gateway with its migrated mTLS config. The secure Snap gateway uses
https://127.0.0.1:17670and requires a client bundle in the user's Snap state. Refresh replaces insecure configs without keeping a copy; follow the published Snap installation steps to re-register an old HTTP client. Unauthenticatedfrom an edge or OIDC gateway: refresh stored credentials withopenshell gateway login [name], then retry. Usegateway logoutonly when intentionally clearing local credentials.- A direct development endpoint with a private or self-signed certificate can be isolated with
--gateway-endpoint <url> --gateway-insecure; do not persist or recommend insecure verification for shared gateways.
Step 2: Identify the Compute Platform
Use gateway metadata, deployment values, or the user's setup notes to identify the driver.
| Platform | Primary checks |
|---|---|
| Docker | Gateway process logs, Docker daemon health, sandbox containers, image pulls. |
| Podman | Podman socket, rootless networking, sandbox containers, image pulls. |
| Kubernetes | Helm release, gateway workload, service, secrets, sandbox pods, events. |
| OpenShift | Same as Kubernetes, plus SecurityContextConstraints (SCCs) and, for external access, an OpenShift Route. Detect OpenShift by the presence of the route.openshift.io API group (oc api-resources --api-group=route.openshift.io). |
| VM | VM driver logs, rootfs availability, host virtualization support. |
| Extension | External driver process, Unix socket ownership/mode, configured driver name, capability handshake, gateway logs. |
Step 3: Check Gateway Startup Dependencies
Before debugging the compute platform, inspect gateway logs for failures in dependencies initialized before the listener becomes ready.
For resource-admission failures, distinguish disabled caller driver config from
missing resource approval. Helm defaults server.drivers.kubernetes.allowDriverConfig
to false and resourceAdmission.enabled to true. Existing PVCs, RuntimeClasses,
and PriorityClasses need matching administrator-owned labels; namespace
membership and read-only access do not grant approval. GPU devices and
operator-selected image-pull Secrets do not need admission labels. In managed
mode, inspect the configured source image-pull Secret in the gateway namespace
and the generation copies (openshell.ai/component=image-pull) in the workspace
namespace. Legacy workloads without
admission provenance need recreation. Do not
automatically label control-plane resources or disable enforcement as a repair.
For out-of-tree compute drivers, also check that their versioned admission-policy
acknowledgement matches the gateway's policy. Configure standalone driver policy
through its administrator-owned --admission-config-json option.
For out-of-tree compute drivers, confirm the selected driver name and socket agree across CLI flags or gateway.toml, and that the operator-owned driver is running before the gateway starts:
rg -n '^version|compute_driver|socket_path|guest_tls_' /etc/openshell/gateway.toml
stat /run/openshell/<driver>.sock
journalctl -u <driver-service> --no-pager --lines=200
journalctl -u openshell-gateway --no-pager --lines=200
Gateway configuration requires [openshell] version = 2, a singular
compute_driver selector, and driver-owned settings under
[openshell.drivers.<name>]. The gateway rejects legacy compute_drivers and
--drivers selectors rather than silently migrating them. One valid, nonempty
OPENSHELL_DRIVERS value remains a deprecated environment-only alias when the
canonical selector is absent; the gateway selects that driver with a warning.
Empty, invalid, comma-delimited, or conflicting values fail startup. The
WebSocket tunnel for edge-proxy CLI access is off unless
enable_websocket_tunnel = true (server.enableWebsocketTunnel in Helm). Homebrew
and RPM package startup migrates only exact package-generated v1 defaults. If
an upgraded package still reports an unsupported version, inspect the active prefix or ~/.config/openshell/gateway.toml; an edited v1 file must
follow the published schema-v2 migration steps and must not be overwritten.
Guest TLS CA, certificate, and key paths are the exception to driver ownership:
configure the complete bundle under [openshell.gateway], and the gateway
injects it only into the selected local driver. TLS-enabled Docker, Podman, and
VM drivers fail startup when neither those paths nor the package-managed local
bundle is available; Kubernetes projects its bundle through a Secret.
Custom names use [openshell.drivers.<name>].socket_path. A launch-time --compute-driver-socket override may also use docker, podman, kubernetes, or vm; the endpoint then takes precedence over built-in construction. First-party standalone drivers require the socket parent directory to be owned by the driver's effective UID, force its mode to 0700, create the socket with mode 0600, and accept only peers with that same UID. Check the parent and socket separately with stat; a gateway running under a different UID cannot connect even when filesystem permissions or group membership would otherwise allow it. Operator-supplied drivers must provide equivalent access control appropriate to their implementation. Check gateway logs for connection errors, GetCapabilities failures, missing peer metadata, protocol-major mismatch, unmet required capabilities, or an unexpected advertised driver name. openshell gateway info reports successful startup negotiations. The advertised name is diagnostic metadata; negotiated features control optional behavior. The gateway does not create or supervise operator-supplied driver processes or sockets.
For the Kubernetes Secrets credential driver, every provider credential lives in
the configured namespace, in every workspace mode. A PermissionDenied error
naming another namespace means the provider's credential handle points outside
the configured namespace; recreate the provider. An unknown field startup
error for [openshell.credential_drivers.kubernetes-secrets] means the table
sets a key the driver does not accept. Confirm the gateway can reach the
credential namespace:
kubectl -n openshell get configmap openshell-config -o jsonpath='{.data.gateway\.toml}' | grep -A3 '^\[openshell\.credential_drivers\.kubernetes-secrets\]'
kubectl auth can-i get secrets -n <credential-namespace> --as system:serviceaccount:openshell:openshell
For a configured Vault credential driver, inspect its endpoint and trust bundle
before debugging provider resolution. Non-loopback addresses must use HTTPS,
and the driver never follows redirects. A private CA bundle augments platform
roots but does not disable hostname verification. With Helm,
server.credentialDrivers.vault.caConfigMapName names a ConfigMap whose
ca.crt key is mounted at /etc/openshell-tls/vault-ca/ca.crt:
kubectl -n openshell get configmap openshell-config -o jsonpath='{.data.gateway\.toml}' | grep -A10 '^\[openshell\.credential_drivers\.vault\]'
kubectl -n openshell get pod -l app.kubernetes.io/name=openshell -o jsonpath='{range .items[0].spec.containers[0].volumeMounts[*]}{.name}{" "}{.mountPath}{"\n"}{end}' | grep vault-ca
kubectl -n openshell get configmap <vault-ca-configmap> -o jsonpath='{.data.ca\.crt}' | openssl x509 -noout -subject -issuer -dates
kubectl -n openshell logs statefulset/openshell -c openshell-gateway --tail=200
An HTTP service DNS address fails configuration validation. UnknownIssuer or
an invalid CA error means the ConfigMap is missing, the ca.crt key is wrong,
or the bundle does not contain the Vault server's issuer. A hostname mismatch
means the HTTPS address host is absent from the server certificate SANs; keep
verification enabled and issue a certificate for the service DNS name.
For configured gateway interceptors, inspect [[openshell.gateway.interceptors]], their Unix or network endpoints, and gateway startup logs:
rg -n 'interceptors|provider_profile_sources|grpc_endpoint|tls_ca_cert_path|audience|allow_insecure_transport|binding_policy|failure_policy|gateway_jwt' /etc/openshell/gateway.toml
stat /run/openshell/interceptors/<name>.sock
journalctl -u <interceptor-service> --no-pager --lines=200
journalctl -u openshell-gateway --no-pager --lines=200
The gateway calls each interceptor's Describe RPC and validates its manifest at startup. Check for missing peer metadata, protocol-major mismatch, unmet required capabilities, unreachable endpoints, invalid RPC/phase bindings, strict allowlist or exact mismatches, and post_commit bindings that resolve to fail_closed. If gateway JWT signing is enabled, authenticated network interceptors require HTTPS and a valid bearer token; check the private CA path, endpoint hostname, expected audience, issuer, kid, and interceptor logs for token rejection. allow_insecure_transport = true explicitly preserves unauthenticated plaintext behavior. If provider_profile_sources names an interceptor, that interceptor must advertise provider-profile capability and return a valid, duplicate-free catalog. A selected interceptor-only source is authoritative; include a user source explicitly when composition is intended. The builtin source type was removed: a config that still names it is rejected at startup.
If the deployment uses supervisor middleware, follow the supervisor middleware troubleshooting reference for startup, authentication, policy validation, and HTTP or WebSocket failures.
For network policy validation failures, first distinguish a gateway mutation
rejection from a supervisor runtime rejection. Direct policy updates,
incremental merges and approvals, provider attachments, and provider-profile
fanout are validated against the complete effective policy before persistence
when the gateway knows the affected sandbox scope. A FAILED_PRECONDITION
ambiguity response means no invalid revision or partial fanout was stored.
Supervisor validation remains defense in depth for startup, races, and policy
sources outside those mutation paths.
Runtime rejection behavior is configured only in gateway.toml:
[openshell.gateway]
policy_validation_failure_mode = "fail_closed"
The default fail_closed mode deactivates the previous generation, closes
pinned relays, and quarantines new egress until a valid generation loads.
retain_last_valid explicitly keeps the previous valid policy active; without
one it still fails closed. Restart the gateway after changing this field.
Inspect sandbox OCSF configuration and finding events for the validation
rationale, configured and effective modes, active generation, and the explicit
previous_policy_active state.
The published supervisor image uses a shell-free distroless Debian 13 base.
Use container logs, engine inspection and the configured exec health probe for
diagnostics; exec ... sh, package installation and in-container shell scripts
are unavailable. Workload shells belong to the separate sandbox image. Preserve
the driver-selected UID and writable runtime/log mounts when reproducing a
supervisor startup failure.
A ConfigurationInvalid readiness condition means startup admission rejected
the image/effective policy or provider configuration. The supervisor remains
alive while the workload stays unstarted. Inspect openshell sandbox get and
repair the desired configuration with a complete policy replacement or provider
change; do not treat a healthy container as proof that the workload is ready.
If the 300-second provisioning repair window expires, the gateway records
ProvisioningTimedOut and stops workload and supervisor compute. Inspect
provisioning in sandbox JSON and TUI NOTES to distinguish cleanup pending from
complete. Repairing configuration after expiry does not restart compute: wait
for cleanup, then explicitly use sandbox start. Repeated rejected reports do
not refresh the deadline, and the CLI wait timeout does not control it.
See policy validation and repair.
The isolated supervisor requests image-policy discovery through the authenticated
sandbox boundary before admission. The workload boundary can remain alive without
launching the workload while configuration is repaired. An unavailable boundary
fails discovery within its control-request deadline. Permanent
gateway errors and exhausted transient retries terminate startup; inspect those
errors as connectivity, authorization, or lifecycle failures.
Step 4: Check Docker-Backed Gateways
docker info
docker ps --filter name=openshell
docker logs <container> --tail=200
docker run --rm --entrypoint /openshell-sandbox "${OPENSHELL_SANDBOX_RUNTIME_IMAGE:-ghcr.io/nvidia/openshell/sandbox:latest}" --version
openshell status
For Docker GPU failures, check CDI support and NVIDIA CDI discovery separately:
docker info --format '{{json .CDISpecDirs}}'
docker info --format '{{json .DiscoveredDevices}}'
for dir in /etc/cdi /var/run/cdi; do
if [ -d "$dir" ]; then
find "$dir" -maxdepth 1 -type f \( -name '*.yaml' -o -name '*.json' \) -print
else
echo "$dir missing"
fi
done
systemctl is-enabled nvidia-cdi-refresh.service nvidia-cdi-refresh.path || true
systemctl is-active nvidia-cdi-refresh.service nvidia-cdi-refresh.path || true
systemctl status nvidia-cdi-refresh.service nvidia-cdi-refresh.path --no-pager --lines=50
journalctl -u nvidia-cdi-refresh.service --no-pager --lines=100
When the NVIDIA Container Toolkit CDI refresh units are not enabled or no NVIDIA CDI spec has been generated, enable them and trigger a refresh:
sudo systemctl enable --now nvidia-cdi-refresh.path
sudo systemctl enable --now nvidia-cdi-refresh.service
sudo systemctl restart nvidia-cdi-refresh.service
docker info --format '{{json .DiscoveredDevices}}'
Common findings:
- Docker daemon unavailable: start Docker Desktop or Docker Engine.
- Gateway process stopped: inspect exit status and logs.
- Sandbox image missing or pull denied: verify image reference and registry credentials.
- Sandbox fails before readiness with an identity-resolution error: inspect the image's OCI
USERand matching/etc/passwdand/etc/groupentries, or explicitly set both process identity fields in policy. Numeric workload identities1through4294967294are accepted; root, the invalid identity sentinel, and missing identities are rejected. - Sandbox fails before readiness with an OCI workspace validation error: inspect the image's
WorkingDirusing the immutable image ID reported by the gateway. Empty,/, and explicit/sandboxuse the managed/sandboxcompatibility workspace. Any other workdir must be an absolute normalized directory with no symlink components; the final policy UID, primary GID, and supplementary groups must pass the kernel's effective traverse/write checks, including POSIX ACL and LSM decisions. OpenShell does not create, chown, or chmod a non-default image workdir. - Docker also rejects an image
VOLUMEthat covers the workdir or one of its parents because the runtime would mask the immutable path before validation. Move theVOLUMEbelow the workspace or remove the declaration. - A workdir rejected as a special filesystem or OpenShell control-path collision cannot be made valid with permissions. Move the image workdir away from kernel-backed mounts and the concrete supervisor, TLS, token, runtime, and socket paths named in the error.
- Local Docker gateway setup cannot copy
openshell-sandboxafter exporting a supervisor image: the sandbox runtime and supervisor are separate artifacts. The runtime image must provide/openshell-sandbox; the supervisor image provides/openshell-supervisor. - Docker driver cannot initialize because it cannot find
openshell-sandbox: verify the sibling binary next toopenshell-gateway, or that the configuredsandbox_runtime_imagecontains/openshell-sandbox. - Sandbox never registers: check gateway logs and the supervisor's gateway endpoint.
- Calls to an external tool server fail while the sandbox is Ready: inspect
Tool server connectionsinopenshell sandbox get <name>. For configured MCP-over-HTTP endpoints, JSON output exposes each address together withlast_resultandlast_reported_atinendpoint_statuses. Select the endpoint by host, path, and ports, then check the reported failure boundary.last_reported_atrecords gateway acceptance time and can advance when retained evidence is accepted after a reset. Results do not expire or prove current availability;HttpResponseReceivedcan still contain a tool error. If several paths share a host and port, a failure before the path is known remains in logs. Verify the actual operation when current tool availability matters. - On Docker Desktop, repeated
Policy fetch failed after 5 attemptsmessages can mean host networking is disabled. Enable host networking in Docker Desktop, ensure Enhanced Container Isolation is disabled, and verify the gateway's primary endpoint is reachable from a host-networked container. - Sandbox runtime image exits before printing
openshell-sandbox --version: verify the configured image contains a static executable at/openshell-sandbox. - A sandbox with explicit
protocol: tcpendpoints fails before workload readiness: confirm the selected isolation backend advertises TCP mediation, then inspect the sandbox and supervisor logs for protected-channel setup or listener failures. A driver that cannot supply the required outer egress fence and authenticated runtime channel must reject the policy before starting the agent. - Supervisor runtime validation fails: verify
supervisor_imagecontains an/openshell-supervisorexecutable from the same release as the sandbox runtime, and that the dynamic loader and shared libraries it links against are available inside that image.docker run --rm --network none --entrypoint /openshell-supervisor <supervisor_image> --versionshould print that release; ano such file or directoryerror for a binary that exists means the loader or a library is missing. The supervisor runs from its own image and does not need to be static; only/openshell-sandboxmust be. - The sandbox fails its enforcement probe: inspect the sandbox log for the exact nested seccomp user-notification, task-memory, Landlock, loopback DNS, or socket-injection check that failed. A runtime may return
ENOSYSforprocess_vm_readvandprocess_vm_writevwhile satisfying the production parent-to-workload-child task-memory probe through/proc/<pid>/mem; only failure of both backends is fatal. Do not add capabilities or switch to an unconfined seccomp profile; use a runtime whose default profile permits the unprivileged probe. - A GPU sandbox fails because Docker reports no discovered NVIDIA CDI devices: verify
.DiscoveredDevicescontains entries such asnvidia.com/gpu=all, verify/etc/cdior/var/run/cdicontains a generated NVIDIA spec, and check thatnvidia-cdi-refresh.serviceandnvidia-cdi-refresh.pathfrom NVIDIA Container Toolkit are enabled and healthy. The service is a one-shot unit, soinactive (dead)can be normal after a successful run; usesystemctl statusandjournalctlto distinguish success from a skipped or failed refresh. Restartnvidia-cdi-refresh.serviceto regenerate missing or stale CDI specs, then restart or reload Docker and re-checkdocker info.
Corporate upstream proxy
Docker corporate proxy settings are operator-owned fields under
[openshell.drivers.docker]. Confirm the complete proxy table and inspect the
companion supervisor command and logs:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 15k
- Forks
- 2k
- Last commit
- Oct 2026
ahel review
K4info
destructive
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
debug-openshell-cluster- Source
- github.com/nvidia/openshell
Related picks
Skill · docker
The pick for Dockerdocker-sandbox
Skill · joelhooks
The pick for Dockerazure-kubernetes
Skill · microsoft
The pick for Kubernetesinfra-containers-kubernetes
Skill · agents-inc
The pick for Kubernetesfind-skills
Skill · vercel-labs
More in Cloud & infravercel-react-best-practices
Skill · vercel-labs
More in Cloud & infra