AlpacaEval
An automatic evaluator for instruction-following models that uses an LLM judge to compare outputs against reference responses, with published leaderboards and length-controlled scoring.
Preview of the business's website.
Services
- LLM-as-judge comparison
- Instruction-following leaderboard
- Length-controlled metric
- Reproducible harness
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.