datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}P3
Dataset Card for P3
Dataset Summary
P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.xP3xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.xP3allxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.xP3mtxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.crows_pairs_multilingualThis is a revised version of CrowS-Pairs that measures stereotypes in language modelling in both English and French.xP3megds
Dataset Card for xP3
Dataset Summary
xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.
Creation: The dataset can be recreated using instructions available here. We provide this version to save processing time and ease reproducibility.
Languages: 46 (Can… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/xP3megds.details_bigscience__bloom-7b1
Dataset Card for Evaluation run of bigscience/bloom-7b1
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-7b1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-7b1.peacock-data-public-datasets-idc-bigscience
bigscience
Research workshop on large language models - The Summer of Language Models 21
At the moment we have 2 code repos:
https://github.com/bigscience-workshop/Megatron-DeepSpeed - this is our flagship code base
https://github.com/bigscience-workshop/bigscience - (this repo) for everything else - docs, experiments, etc.
Currently, the most active segments of this repo are:
JZ - Lots of information about our work environment which helps evaluate, plan and get things done… See the full description on the dataset page: https://huggingface.co/datasets/applied-ai-018/peacock-data-public-datasets-idc-bigscience.shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.bigscience-lama
Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models.
@inproceedings{petroni2020how,
title={How Context Affects Language Models' Factual Predictions},
author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel},
booktitle={Automated Knowledge Base Construction},
year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters"
More Information needed
olm-october-2022-tokenized-1024-no-bigscience-filters
Dataset Card for "olm-october-2022-tokenized-1024-no-bigscience-filters"
More Information needed
details_bigscience__bloom
Dataset Card for Evaluation run of None
Dataset Summary
Dataset automatically created during the evaluation run of model None on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom.lm_code_github-eval_subsetcollaborative_catalogdetails_bigscience__bloom-560m
Dataset Card for Evaluation run of bigscience/bloom-560m
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-560m on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 13 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-560m.HIPE2020_sent-splitTODOhipe2020TODOdetails_bigscience__bloom-1b1
Dataset Card for Evaluation run of bigscience/bloom-1b1
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-1b1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-1b1.roots_en_wikipediaROOTS Subset: roots_en_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.Open_Medieval_French
Open Medieval French
Source: https://github.com/OpenMedFr/texts
details_bigscience__bloomz-7b1
Dataset Card for Evaluation run of bigscience/bloomz-7b1
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloomz-7b1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloomz-7b1.roots_en_no_code_stackexchangeROOTS Subset: roots_en_no_code_stackexchange
Stack Exchange Website
Dataset uid: no_code_stackexchange
Description
Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers.
Homepage
https://stackexchange.com/
Licensing
open license
cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.tokenizer-probing-ud2.10bigscience__bloom-7b1-details
Dataset Card for Evaluation run of bigscience/bloom-7b1
Dataset automatically created during the evaluation run of model bigscience/bloom-7b1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-7b1-details.details_bigscience__bloom-1b7
Dataset Card for Evaluation run of bigscience/bloom-1b7
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloom-1b7 on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-1b7.details_bigscience__bloomz-3b
Dataset Card for Evaluation run of bigscience/bloomz-3b
Dataset Summary
Dataset automatically created during the evaluation run of model bigscience/bloomz-3b on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloomz-3b.bigscience__bloom-3b-details
Dataset Card for Evaluation run of bigscience/bloom-3b
Dataset automatically created during the evaluation run of model bigscience/bloom-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-3b-details.bigscience__bloom-560m-details
Dataset Card for Evaluation run of bigscience/bloom-560m
Dataset automatically created during the evaluation run of model bigscience/bloom-560m
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-560m-details.
