ahnafyy/skills-evals

Validate, trigger-test, and regression-test agent skills, instructions, and custom agents (Copilot, Claude, Cursor)

View on GitHub

Trust Signals

Scorecard Score
not yet scored
Maintenance Recency
Stale
License
None
namedescriptionrequireddefault
rootDirectory to scan for agent artifactsno.
argsExtra CLI arguments passed to `skills-evals run`no""
behavioralComma- or space-separated artifact names to run behavioral (Tier 3) evals for. Requires the executor adapter CLI to be installed and authenticated in the job. Best run on a schedule against a cheap model. Empty (default) skips behavioral evals entirely.no""
adapterBehavioral executor adapter: claude, copilot, or cursor. Empty falls back to behavioral.adapter in skills-evals.config.json, then claude.no""
graderBehavioral grader adapter. Empty falls back to config, then the executor adapter.no""

no outputs