datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
P3
Dataset Card for P3
Dataset Summary
P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.BIGstockimage-1.5MBIGstockimage-1.5M-scored-pt-twoBIGstockimage-1.5M-scored-pt-oneopenai_MMMLU_zhoolm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters"
More Information needed
dclm-baseline-subsetopenai_MMMLU_engopenai_MMMLU_arbopenai_MMMLU_hindynasample_trainbigsurvey_with_sent_srl_scoresdynasample_multitasks_cleanopenai_MMMLU_spadynasample_train_scoreby3llmsopenai_MMMLU_rusopenai_MMMLU_swaroots_en_wikipediaROOTS Subset: roots_en_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.Open_Medieval_French
Open Medieval French
Source: https://github.com/OpenMedFr/texts
roots_en_no_code_stackexchangeROOTS Subset: roots_en_no_code_stackexchange
Stack Exchange Website
Dataset uid: no_code_stackexchange
Description
Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers.
Homepage
https://stackexchange.com/
Licensing
open license
cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.bigsurvey_with_srlopenai_MMMLU_deuroots_indic-bn_wikipediaROOTS Subset: roots_indic-bn_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-bn_wikipedia.roots_zh-cn_wikipediaROOTS Subset: roots_zh-cn_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-cn_wikipedia.ultralinkbigsurvey_with_srl_newroots_es_wikipediaROOTS Subset: roots_es_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_wikipedia.roots_en_book_dash_booksROOTS Subset: roots_en_book_dash_books
Book Dash Books
Dataset uid: book_dash_books
Description
Book Dash believes that every child should own one hundred books by the age of five.
To that end, we gather creative professionals who volunteer to create new, African storybooks that anyone can freely translate, print and distribute. In this way, we have vastly reduced the costs involved in putting high-quality books in children’s hands and hearts.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_book_dash_books.bi-gsm8k
Bi-GSM8K
A high-quality bilingual dataset of 500 elementary-level math problems with step-by-step correct solutions and annotated student error patterns in both English and Korean.
📋 Overview
Bi-GSM8K is designed to support intelligent tutoring systems by providing:
Teacher-authored correct solutions with step-level details
Expert-created simulated student solutions reflecting common mathematical misconceptions actually observed in real students
Explicit step-level… See the full description on the dataset page: https://huggingface.co/datasets/Tutoruslabs/bi-gsm8k.roots_fr_wikipediaROOTS Subset: roots_fr_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_fr_wikipedia.
