dileepkpandiya/skilleval

Measure whether a Claude SKILL.md file improves model outputs using a blind A/B test and LLM judge.

View on GitHub

Trust Signals

Scorecard Score
not yet scored
Maintenance Recency
Stale
License
None
namedescriptionrequireddefault
skill-pathPath to directory containing SKILL.mdyes
tasksPath to tasks YAML fileno""
fail-belowFail if skill effectiveness is below this thresholdno""
fail-if-hurt-pctFail if percentage of hurt tasks exceeds this valueno""
judge-providerJudge provider - gemini-flash | claude | openainogemini-flash
runsNumber of independent A/B runs per taskno1
anthropic-api-keyAnthropic API key for the runneryes
gemini-api-keyGemini API key (required if judge-provider is gemini-flash)no""
openai-api-keyOpenAI API key (required if judge-provider is openai)no""
namedescription
avg-diffAverage effectiveness score diff
tasks-improvedNumber of tasks improved
tasks-hurtNumber of tasks hurt
confidenceOverall confidence rating