Evaluation
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}NexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.GDP-Val-Evaluation-Submission
GDPval Submission Dataset
This dataset contains model outputs for GDP-Val evaluation.
Dataset Structure
data/: Contains the main dataset in Parquet format
train-00000-of-00001.parquet: Submission data with model outputs
deliverable_files/: Contains generated files for tasks that produce file deliverables
Organized by task_id
dataset_info.json: Metadata about the dataset
Columns
task_id: Unique identifier for each task
sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.grobid-evaluation
GROBID End-to-End Evaluation Dataset
Reference corpora used for GROBID end-to-end
benchmarking of scientific-article structuring.
Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/
Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/
Official archive (Zenodo): https://zenodo.org/record/7708580
Dataset summary
These are the datasets used for GROBID end-to-end benchmarking, covering:
metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.da-code-evaluation-resultsaya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.
