datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
auto-benchmarkcards
Auto-Generated BenchmarkCards
A catalog of structured documentation cards for AI evaluation benchmarks. Each card is an LLM-composed, source-grounded summary of a benchmark's purpose, data, methodology, risks, and limitations. The dataset exists to make benchmark documentation consistent, comparable, and easy to inspect across tasks and domains.
Dataset Details
Language(s): English
License: Community Data License Agreement, Permissive, Version 2.0
Schema: based… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/auto-benchmarkcards.BenchmarkCards
Dataset Card for BenchmarkCards
BenchmarkCards is a standardized documentation dataset for large language model (LLM) benchmarks.
Each card summarizes key information about an LLM benchmark, including its objectives, methodology, data sources, targeted risks, limitations, and ethical considerations.
🙏 Acknowledgments
We gratefully thank all benchmark authors who provided feedback and approval for the BenchmarkCards in this repository. Your collaboration is essential… See the full description on the dataset page: https://huggingface.co/datasets/ASokol/BenchmarkCards.
