datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
helm-scenarios
HELM Scenarios
This repository contains mirrors of datasets that are used as scenarios by crfm-helm.
Scenarios
TURL Column Type Annotation
The subfolder turl-column-type-annotation contains files for the table column type annotation task from the TURL paper. No modifications were made to these files.
The TURL dataset by Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu is licensed under CC BY 4.0. The TURL dataset was modified from the TabEL… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/helm-scenarios.train_splits_helmContains the following train split from datasets in helm:
big bench
mmlu
TruthfulQA
cnn/dm
gsm
bbq
boolq
NarrativeQA
QuAC
math
bAbI
Each prompt has <= 5 in-context samples along with a sample, all of which from the train set of the respective datasets.
HELMET
HELMET: How to Evaluate Long-context Language Models Effectively and Thoroughly
[Paper][Code]
HELMET is a comprehensive benchmark for long-context language models covering seven diverse categories of tasks.
The datasets are application-centric and are designed to evaluate models at different lengths and levels of complexity.
Please check out the paper for more details, and the code repo for how to process the data and run the evaluations
quac_helmnarrative_qa_helmcodellava-helmet-plaincodellava-helmet-memwrap
