vishisht16/HumaneProxy Safety Benchmark
Run AI safety evaluations through HumaneProxy's pipeline to catch prompt regressions before production.
View on GitHubTrust Signals
- Scorecard Score
- not yet scored
- Maintenance Recency
- Stale
- License
- None
Inputs
| name | description | required | default |
|---|---|---|---|
| dataset | Path to a JSON evaluation dataset. Each entry needs 'message' and 'expected' (safe | self_harm | criminal_intent). | yes | — |
| python-version | Python version to use. | no | 3.12 |
| extra | pip install extras. Comma-separated. Defaults to 'ml' because `hp benchmark` runs stages 1+2 by default and Stage 2 silently no-ops without the sentence-transformers dependency. Pass '' for heuristics-only. | no | ml |
Outputs
no outputs