invincible-jha/AumOS Agent Evaluation

Run agent evaluations and benchmarks in your CI pipeline with multi-run statistical analysis

View on GitHub

Trust Signals

Scorecard Score
not yet scored
Maintenance Recency
Stale
License
None
namedescriptionrequireddefault
benchmark-pathPath to benchmark YAML fileyes
eval-runsNumber of evaluation runs for statistical significanceno5
python-versionPython versionno3.12
fail-thresholdMinimum score threshold to pass (0.0-1.0)no0.7
output-formatReport format: json, markdown, htmlnomarkdown
dimensionsComma-separated evaluation dimensions: accuracy,safety,consistency,costnoaccuracy,safety,consistency,cost
namedescription
overall-scoreOverall evaluation score
report-pathPath to evaluation report
passedWhether the evaluation passed the threshold