Stanford HELM
A holistic evaluation framework from Stanford CRFM that scores models across many scenarios and metrics including accuracy, calibration, robustness and bias, and publishes public leaderboards.
Preview of the business's website.
Services
- Multi-metric evaluation
- Scenario library
- Public leaderboards
- Reproducible runs
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.