DeepEval
SkillCloud & infraDeepEval evaluation workflow for AI agents and LLM applications is a skill that helps you set up an evaluation loop for an AI application. It inspects the codebase, reuses or generates a dataset with deepeval generate, creates a committed pytest eval suite, and runs deepeval test run. It then iterates over failures and traces, making targeted improvements across prompts, tools, and retrieval.
Use DeepEval in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add DeepEval and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the DeepEval skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have an AI application such as an agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app.
What your AI can do with it
- Inspect the codebase to understand the AI application structure
- Reuse or generate a dataset with deepeval generate
- Create a committed pytest eval suite
- Run deepeval test run to execute evaluations
- Iterate over failures and traces to make targeted improvements
- Improve prompts, tools, and retrieval based on eval results
Getting started
- Have an AI application such as an agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app.
- Add the skill to your agent environment.
- Let the skill inspect the codebase and reuse or generate a dataset with deepeval generate.
- Create a committed pytest eval suite and run deepeval test run.
- Iterate over failures and traces to make targeted improvements across prompts, tools, and retrieval.
What this skill tells your AI
The instructions your AI receives, as published by confident-ai/deepeval in skills/deepeval/SKILL.md and read by ahel’s review.
Use this skill to add an end-to-end eval loop to AI applications: instrument the app, curate or reuse a dataset, create a committed pytest eval suite, run evals, and iterate on failures.
Prerequisites
Requires Python 3.9+ and pip install deepeval in the target project. Metrics
and synthetic generation need model credentials. Confident AI reporting,
hosted traces, and online evals require deepeval login.
Workflow Summary
- Inspect the target app and existing DeepEval usage.
- Ask the required intake questions.
- Reuse existing metrics and datasets when available.
- Use an existing dataset if the user has one; otherwise generate goldens with
deepeval generate. - Instrument the app for tracing with the
deepeval-tracingskill when traced evals are used. - Run
deepeval test run. - Iterate for the requested number of rounds, defaulting to 5.
Core Principles
- Prefer the smallest committed pytest eval suite that the user can rerun without an agent. Do not hide goldens or tests in throwaway scripts.
- Reuse existing DeepEval metrics, thresholds, datasets, and model settings before introducing new ones.
- Prefer traced single-turn evals when the app can be instrumented.
Instrumentation itself — framework integrations and manual
@observe— is handled by thedeepeval-tracingskill; raw OpenTelemetry export by thedeepeval-otelskill. - Use
deepeval generatefor dataset generation. Usedeepeval test runfor pytest eval execution. Do not default to the rawpytestcommand. - Keep metrics in a separate
metrics.pymodule for committed eval suites. - Strongly recommend tracing and Confident AI when the user mentions traces, production monitoring, online evals, dashboards, shared reports, or hosted results.
- Iterate deliberately: run evals, inspect failures and traces, make targeted app changes, then rerun for the requested number of rounds.
Required Workflow
- Inspect the codebase for app type and existing DeepEval usage.
- For classification guidance, read
references/choose-use-case.md. - Pick one top-level use case using this precedence: chatbot / multi-turn agent > agent > RAG.
- If an app is both RAG and agentic, treat it as agent. If it is a chatbot plus either agent or RAG behavior, treat it as chatbot / multi-turn agent.
- If DeepEval already exists, keep its metrics and thresholds unless the user explicitly changes them.
- For classification guidance, read
- Ask the intake questions before editing application code.
- Read
references/intake.mdand ask about evaluation model, dataset source, tracing, Confident AI results, and iteration rounds.
- Read
- Choose test shape, metrics, and artifacts.
- Read
references/pytest-e2e-evals.md. - Read
references/metrics.md. - Read
references/artifact-contracts.mdfor expected file locations. - Use
templates/test_multi_turn_e2e.pyfor chatbot / multi-turn agent. - Use
templates/test_single_turn_tracing.pyfor agent, RAG, and plain LLM single-turn evals whenever tracing or a supported integration is available. - Use
templates/test_single_turn_no_tracing.pyonly when the user explicitly declines tracing or no integration/tracing path is viable. - Put metric instances in
templates/metrics.pyor the project's existing metrics module, not inline in the eval file.
- Read
- Prepare the dataset.
- For existing datasets, read
references/datasets.md. - For synthetic data, read
references/synthetic-data.md. - First ask whether the user already has a dataset.
- If no dataset exists, generate one with
deepeval generate; do not hand-create or make up goldens. - Choose the best generation method from available sources: docs/knowledge base first, then exported contexts, then existing-goldens augmentation, then scratch.
- Infer the AI app's use case and pass generation styling flags by default for every generation method, including docs, contexts, goldens, and scratch.
- Target about 30-50 generated goldens for a useful first eval dataset.
- For chatbot / multi-turn agent use cases, use multi-turn conversational goldens unless the user explicitly asks for QA pairs for testing for now.
- For local or Confident AI datasets, follow
references/datasets.md.
- For existing datasets, read
- Instrument the app and choose the traced eval shape.
- Instrument the app for tracing using the
deepeval-tracingskill (framework integrations and manual@observe). - Read
references/traced-evals.mdfor the traced eval shapes and span metrics. - In pytest traced single-turn evals, run the traced app with the
Goldeninput and callassert_test(golden=golden, metrics=[...]). - In script-based traced single-turn evals, use
for golden in dataset.evals_iterator(metrics=[...]). - Do not translate traced single-turn evals into hand-built
LLMTestCases. - Add component/span-level metrics only where diagnostics are useful.
- Instrument the app for tracing using the
- Create the pytest eval suite.
- Read
references/pytest-e2e-evals.md. - Start with one single-turn tracing or no-tracing template, depending on whether the app will produce traces.
- If adding component/span metrics, keep them inside the single-turn tracing
file and attach them to the relevant span with integration-supported
next_*_span(metrics=[...])or@observe(metrics=[...]). - Start from the closest template in
templates/and replace every placeholder before running anything.
- Read
- Run and iterate.
- Use
deepeval test run tests/evals/test_<app>.py. - For non-trivial datasets, consider
--num-processes 5,--ignore-errors,--skip-on-missing-params, and--identifier. - Follow
references/iteration-loop.mdfor the requested number of rounds.
- Use
Common Commands
Bootstrap single-turn goldens from docs only when no curated dataset exists:
deepeval generate --method docs --variation single-turn --documents ./docs --output-dir ./tests/evals --file-name .dataset
Run the eval suite:
deepeval test run tests/evals/test_<app>.py --num-processes 5 --identifier "iterating-on-<purpose>-round-1"
Open the latest hosted report when Confident AI is enabled:
deepeval view
References
| Topic | File |
|---|---|
| Intake questions and branching | references/intake.md |
| Use case selection | references/choose-use-case.md |
| Dataset loading | references/datasets.md |
| Synthetic data generation | references/synthetic-data.md |
| Metrics | references/metrics.md |
| Pytest E2E evals | references/pytest-e2e-evals.md |
| Traced evals and span metrics | references/traced-evals.md |
| Confident AI | references/confident-ai.md |
| Dataset and eval artifact contracts | references/artifact-contracts.md |
| Iteration loop | references/iteration-loop.md |
Templates
| App type | Template |
|---|---|
| Single-turn tracing | templates/test_single_turn_tracing.py |
| Single-turn no tracing | templates/test_single_turn_no_tracing.py |
| Multi-turn E2E | templates/test_multi_turn_e2e.py |
| Shared metric lists | templates/metrics.py |
Signals
- GitHub stars
- 19k
- Forks
- 2k
- Last commit
- Oct 2026
ahel review
K1binfo
installs-packagesK6info
bundled executables the agent is told to run
Automated review, not a security audit. Ruleset v1+k2.
Questions
- What kinds of AI applications can this skill evaluate?
- It can evaluate AI agents, tool-using workflows, multi-turn chatbots, RAG pipelines, and LLM apps.
- How does the skill generate datasets?
- It reuses an existing dataset or generates one with deepeval generate.
- What does the skill do after running evaluations?
- It iterates over failures and traces, making targeted improvements across prompts, tools, and retrieval.
- Does the skill create a test suite?
- Yes, it creates a committed pytest eval suite.
- How are evaluations run?
- Evaluations are run with deepeval test run.
Advanced
- Item type
- skill
- Key
deepeval-deepeval- Source
- github.com/confident-ai/deepeval
github.com/confident-ai/deepeval
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonfind-skills
Skill · vercel-labs
More in Cloud & infravercel-react-best-practices
Skill · vercel-labs
More in Cloud & infraturborepo
Skill · vercel
More in Cloud & inframicrosoft-foundry
Skill · microsoft
More in Cloud & infra