Eval Runner

SubagentDatabases & data

LLM evaluation specialist who runs structured eval datasets, computes quality metrics using DeepEval/RAGAS, tracks regression across model versions, and reports to Langfuse for tracing and scoring.

Delivery for this kind is on the roadmap — not serving yet. You can still add it. It stays paused until ahel can serve it.

Serve it through your gateway

One link, every agent. Your own credentials, stored once.

Signals

GitHub stars
224
Forks
23
Last commit
Aug 2026
Installs
197 stars