alepot55/Agentrial CI

Statistical evaluation for AI agents - test your agents with confidence intervals and failure attribution

View on GitHub

Trust Signals

Scorecard Score
not yet scored
Maintenance Recency
Stale
License
None
namedescriptionrequireddefault
trialsNumber of trials per test caseno10
thresholdMinimum pass rate (0-1) to pass CIno0.9
configPath to agentrial.yml config filenoagentrial.yml
test-pathPath to test files or directorynotests/
python-versionPython version to useno3.11
post-commentPost results as PR commentnotrue
namedescription
pass-rateOverall pass rate
passedWhether the evaluation passed the threshold
results-filePath to JSON results file