Score

SubagentMonitoring & ops

Designs honest model evaluation frameworks — metric selection matched to business cost functions, statistical significance testing, calibration, and confusion analysis. Use when choosing eval metrics or comparing models. Trigger with "design a model evaluation framework", "compare these models stati

Delivery for this kind is on the roadmap — not serving yet. You can still add it. It stays paused until ahel can serve it.

Serve it through your gateway

One link, every agent. Your own credentials, stored once.

Signals

GitHub stars
3k
Forks
392
Last commit
Aug 2026
Installs
2k stars