secure-ai-agent-coding

SkillMonitoring & ops

Build, review, or harden AI agents, LLM apps, RAG, tool-calling, coding agents, and agentic workflows with secure-by-default controls. Trigger on prompt injection, tool permissions, approvals, AI data handling, RAG poisoning, observability, or safety CI. Do NOT use for non-AI appsec.

Use secure-ai-agent-coding in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add secure-ai-agent-coding and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the secure-ai-agent-coding skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

secure-ai-agent-codingStart free

What this skill tells your AI

The instructions your AI receives, as published by jpcaparas/skills in skills/agents/secure-ai-agent-coding/SKILL.md and read by Ahel’s review.

Build or review AI agents and LLM applications so their blast radius stays small, their behavior is auditable, and high-impact actions require explicit control.

Scale controls to the actual trust boundaries, data sensitivity, effects, and deployment. A read-only public-data assistant, a disposable coding sandbox, and a production payment agent need different controls. Keep authorization and isolation outside the model; leave architecture and implementation choices open when they preserve those guarantees. The reference catalog is a risk checklist, not a requirement to add every mechanism to every feature.

Decision Tree

What is the user asking for?

  • Build a new AI agent or LLM feature: Use references/implementation-patterns.md for unfamiliar control boundaries and references/controls.md for the relevant risks.
  • Review an existing codebase or design: Use references/review-workflow.md to scope the review. The optional scanner can locate leads in unfamiliar source; known code paths can be inspected directly.
  • Add tools, system calls, code execution, database writes, API calls, or email/message sending: Classify the concrete effect before selecting controls. Read references/implementation-patterns.md for enforced capability boundaries; require per-action authorization and explicit approval or an accepted equivalent for high-impact effects, not repeated human permission for already-authorized sandbox work.
  • Handle user data, production data, documents, web pages, email, vector stores, embeddings, or fine-tuning data: Use references/threat-model.md for injection and data-flow risks; add references/governance.md when production, consent, retention, or lifecycle obligations are involved.
  • Debug a safety incident, unexpected model behavior, prompt injection, data leak, or harmful automation: Read references/review-workflow.md and references/gotchas.md, preserve logs, stop autonomous actions, and recover from a known safe state.
  • The request is only generic web app security with no AI, model, RAG, tool-call, or agentic workflow: Do not use this skill unless the AI-specific attack surface is part of the task.

Quick Reference

TaskAction
Find likely dangerous patternsRun python3 scripts/scan_patterns.py /path/to/project
Classify riskTier each agent action by impact, reversibility, data sensitivity, and external side effects
Protect promptsSeparate instructions from untrusted data; validate, delimit, and minimize context
Protect tool callsUse explicit allowlists, scoped credentials, per-action authorization, and rate limits
Protect users and dataClassify data, minimize disclosure, redact logs, verify consent, and avoid raw production data in test systems
Protect downstream systemsValidate structured model output before using it in code, SQL, shell, APIs, or UI rendering
Protect productionAdd monitoring, anomaly alerts, safety regression tests, rollback plans, and incident response steps
Review exceptionsDocument why a control does not apply, who accepted the risk, and when it expires

Core Workflow

  1. Map the AI surface: model calls, prompts, retrieved content, tools, state, credentials, data stores, and external side effects.
  2. Classify every input as untrusted unless it is generated by a trusted server-side component and still validate it.
  3. Classify every action the agent can take. Default to reversible, low-privilege actions and escalate high-impact actions to human approval.
  4. Select controls at the affected boundaries. Schemas, allowlists, server-side authorization, idempotency, rate limits, locks, rollback, and audit logs solve different risks; use the ones the action and environment require.
  5. Keep sensitive data out of prompts and logs unless the model truly needs it. Prefer redaction, tokenization, anonymization, or synthetic data.
  6. Validate AI output before it reaches downstream interpreters, renderers, databases, APIs, file systems, or shell commands.
  7. Test affected safety and output contracts after model, prompt, framework, dependency, retrieval, or tool changes. Run broader regression coverage when shared boundaries change.

Control Router

AreaLoad
End-to-end review process, severity, report formatreferences/review-workflow.md
Control catalog and evidence checklistreferences/controls.md
Copyable implementation patternsreferences/implementation-patterns.md
AI-specific threat scenarios and risk tiersreferences/threat-model.md
Inventory, consent, model updates, lifecycle, and operationsreferences/governance.md
Common failure modes and reviewer trapsreferences/gotchas.md
Source conversion notes and scope exclusionsreferences/source-policy.md

Operational Scripts

Use scripts/scan_patterns.py as a first-pass heuristic scanner, not as proof of safety.

python3 scripts/scan_patterns.py /path/to/project
python3 scripts/scan_patterns.py /path/to/project --json --fail-on high

The scanner intentionally favors review prompts over automated verdicts. Treat findings as places to inspect.

Source Scope

This skill adapts the engineering-applicable controls from Galdren's "Secure AI & Agent Coding Policy" into an agent skill. It keeps the parts that can drive code review, architecture decisions, local checks, and production hardening. It leaves legal judgment, organization-specific approvals, and physical security as prompts to involve the right team rather than pretending a coding skill can resolve them.

When bundled advice fails or a provider/framework contract changes, consult relevant official version-matched documentation and trusted primary guidance such as OWASP, NIST, or MITRE. Verify the affected control safely; current documentation does not authorize broader access or data disclosure. State unavailable evidence. Propose a canonical skill correction with the affected rule, source/version, and threat scenario or regression, rather than silently updating an installed copy or publishing it.

Gotchas

  1. Prompt-only defenses are not enough. Enforce authorization, validation, and tool limits outside the model.
  2. A model used as a validator is still an AI input surface. Validate its inputs and outputs like any other component.
  3. "Authenticated user" is not the same as "authorized agent action." Check each action separately.
  4. Logs can become a data leak. Record enough to investigate, but redact secrets, personal data, and sensitive context.
  5. Model or framework upgrades can change behavior without a code diff. Retest safety and output contracts after every change.
  6. Broad tool access turns small prompt failures into production incidents. Keep the agent footprint minimal at every step.

Signals

GitHub stars
54
Forks
3
Last commit
Sep 2026

Ahel review

  • K6low
    bundled executables the agent is told to run
  • K3info
    injection (in references/threat-model.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
secure-ai-agent-coding
Source
github.com/jpcaparas/skills