secure-ai-agent-coding
SkillMonitoring & opsBuild, review, or harden AI agents, LLM apps, RAG, tool-calling, coding agents, and agentic workflows with secure-by-default controls. Trigger on prompt injection, tool permissions, approvals, AI data handling, RAG poisoning, observability, or safety CI. Do NOT use for non-AI appsec.
Use secure-ai-agent-coding in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add secure-ai-agent-coding and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the secure-ai-agent-coding skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by jpcaparas/skills in skills/agents/secure-ai-agent-coding/SKILL.md and read by Ahel’s review.
Build or review AI agents and LLM applications so their blast radius stays small, their behavior is auditable, and high-impact actions require explicit control.
Scale controls to the actual trust boundaries, data sensitivity, effects, and deployment. A read-only public-data assistant, a disposable coding sandbox, and a production payment agent need different controls. Keep authorization and isolation outside the model; leave architecture and implementation choices open when they preserve those guarantees. The reference catalog is a risk checklist, not a requirement to add every mechanism to every feature.
Decision Tree
What is the user asking for?
- Build a new AI agent or LLM feature:
Use
references/implementation-patterns.mdfor unfamiliar control boundaries andreferences/controls.mdfor the relevant risks. - Review an existing codebase or design:
Use
references/review-workflow.mdto scope the review. The optional scanner can locate leads in unfamiliar source; known code paths can be inspected directly. - Add tools, system calls, code execution, database writes, API calls, or email/message sending:
Classify the concrete effect before selecting controls. Read
references/implementation-patterns.mdfor enforced capability boundaries; require per-action authorization and explicit approval or an accepted equivalent for high-impact effects, not repeated human permission for already-authorized sandbox work. - Handle user data, production data, documents, web pages, email, vector stores, embeddings, or fine-tuning data:
Use
references/threat-model.mdfor injection and data-flow risks; addreferences/governance.mdwhen production, consent, retention, or lifecycle obligations are involved. - Debug a safety incident, unexpected model behavior, prompt injection, data leak, or harmful automation:
Read
references/review-workflow.mdandreferences/gotchas.md, preserve logs, stop autonomous actions, and recover from a known safe state. - The request is only generic web app security with no AI, model, RAG, tool-call, or agentic workflow: Do not use this skill unless the AI-specific attack surface is part of the task.
Quick Reference
| Task | Action |
|---|---|
| Find likely dangerous patterns | Run python3 scripts/scan_patterns.py /path/to/project |
| Classify risk | Tier each agent action by impact, reversibility, data sensitivity, and external side effects |
| Protect prompts | Separate instructions from untrusted data; validate, delimit, and minimize context |
| Protect tool calls | Use explicit allowlists, scoped credentials, per-action authorization, and rate limits |
| Protect users and data | Classify data, minimize disclosure, redact logs, verify consent, and avoid raw production data in test systems |
| Protect downstream systems | Validate structured model output before using it in code, SQL, shell, APIs, or UI rendering |
| Protect production | Add monitoring, anomaly alerts, safety regression tests, rollback plans, and incident response steps |
| Review exceptions | Document why a control does not apply, who accepted the risk, and when it expires |
Core Workflow
- Map the AI surface: model calls, prompts, retrieved content, tools, state, credentials, data stores, and external side effects.
- Classify every input as untrusted unless it is generated by a trusted server-side component and still validate it.
- Classify every action the agent can take. Default to reversible, low-privilege actions and escalate high-impact actions to human approval.
- Select controls at the affected boundaries. Schemas, allowlists, server-side authorization, idempotency, rate limits, locks, rollback, and audit logs solve different risks; use the ones the action and environment require.
- Keep sensitive data out of prompts and logs unless the model truly needs it. Prefer redaction, tokenization, anonymization, or synthetic data.
- Validate AI output before it reaches downstream interpreters, renderers, databases, APIs, file systems, or shell commands.
- Test affected safety and output contracts after model, prompt, framework, dependency, retrieval, or tool changes. Run broader regression coverage when shared boundaries change.
Control Router
| Area | Load |
|---|---|
| End-to-end review process, severity, report format | references/review-workflow.md |
| Control catalog and evidence checklist | references/controls.md |
| Copyable implementation patterns | references/implementation-patterns.md |
| AI-specific threat scenarios and risk tiers | references/threat-model.md |
| Inventory, consent, model updates, lifecycle, and operations | references/governance.md |
| Common failure modes and reviewer traps | references/gotchas.md |
| Source conversion notes and scope exclusions | references/source-policy.md |
Operational Scripts
Use scripts/scan_patterns.py as a first-pass heuristic scanner, not as proof of safety.
python3 scripts/scan_patterns.py /path/to/project
python3 scripts/scan_patterns.py /path/to/project --json --fail-on high
The scanner intentionally favors review prompts over automated verdicts. Treat findings as places to inspect.
Source Scope
This skill adapts the engineering-applicable controls from Galdren's "Secure AI & Agent Coding Policy" into an agent skill. It keeps the parts that can drive code review, architecture decisions, local checks, and production hardening. It leaves legal judgment, organization-specific approvals, and physical security as prompts to involve the right team rather than pretending a coding skill can resolve them.
When bundled advice fails or a provider/framework contract changes, consult relevant official version-matched documentation and trusted primary guidance such as OWASP, NIST, or MITRE. Verify the affected control safely; current documentation does not authorize broader access or data disclosure. State unavailable evidence. Propose a canonical skill correction with the affected rule, source/version, and threat scenario or regression, rather than silently updating an installed copy or publishing it.
Gotchas
- Prompt-only defenses are not enough. Enforce authorization, validation, and tool limits outside the model.
- A model used as a validator is still an AI input surface. Validate its inputs and outputs like any other component.
- "Authenticated user" is not the same as "authorized agent action." Check each action separately.
- Logs can become a data leak. Record enough to investigate, but redact secrets, personal data, and sensitive context.
- Model or framework upgrades can change behavior without a code diff. Retest safety and output contracts after every change.
- Broad tool access turns small prompt failures into production incidents. Keep the agent footprint minimal at every step.
Signals
- GitHub stars
- 54
- Forks
- 3
- Last commit
- Sep 2026
Ahel review
K6low
bundled executables the agent is told to runK3info
injection (in references/threat-model.md)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
secure-ai-agent-coding- Source
- github.com/jpcaparas/skills
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythoninternal-comms
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opspricing
Skill · coreyhaines31
More in Monitoring & opslark-okr
Skill · larksuite
More in Monitoring & ops