LM Evaluation Harness
EleutherAI's harness for few-shot benchmarking of language models across hundreds of academic tasks. It standardises prompt formatting so results are comparable between models and backends.
Preview of the business's website.
Services
- Hundreds of benchmark tasks
- Few-shot prompt formatting
- Multiple inference backends
- Reproducible scoring
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.