datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
P3
Dataset Card for P3
Dataset Summary
P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters"
More Information needed
olm-october-2022-tokenized-1024-no-bigscience-filters
Dataset Card for "olm-october-2022-tokenized-1024-no-bigscience-filters"
More Information needed
roots_en_wikipediaROOTS Subset: roots_en_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.Open_Medieval_French
Open Medieval French
Source: https://github.com/OpenMedFr/texts
roots_en_no_code_stackexchangeROOTS Subset: roots_en_no_code_stackexchange
Stack Exchange Website
Dataset uid: no_code_stackexchange
Description
Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers.
Homepage
https://stackexchange.com/
Licensing
open license
cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.roots_indic-bn_wikipediaROOTS Subset: roots_indic-bn_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-bn_wikipedia.roots_zh-cn_wikipediaROOTS Subset: roots_zh-cn_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-cn_wikipedia.roots_es_wikipediaROOTS Subset: roots_es_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_wikipedia.roots_en_book_dash_booksROOTS Subset: roots_en_book_dash_books
Book Dash Books
Dataset uid: book_dash_books
Description
Book Dash believes that every child should own one hundred books by the age of five.
To that end, we gather creative professionals who volunteer to create new, African storybooks that anyone can freely translate, print and distribute. In this way, we have vastly reduced the costs involved in putting high-quality books in children’s hands and hearts.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_book_dash_books.roots_en_odiencorpROOTS Subset: roots_en_odiencorp
OdiEnCorp2.0
Dataset uid: odiencorp
Description
OdiEnCorp is a collection of Odia-English parallel and Odia monolingual sentences collected from different sources such as Odia Wikipedia, web sites, books, and dictionaries using different manual and machine learning techniques including web scraping and optical character recognition. OdiEnCorp 2.0 served in WAT 2020 EnglishOdia Indic Task.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_odiencorp.roots_indic-or_odiencorpROOTS Subset: roots_indic-or_odiencorp
OdiEnCorp2.0
Dataset uid: odiencorp
Description
OdiEnCorp is a collection of Odia-English parallel and Odia monolingual sentences collected from different sources such as Odia Wikipedia, web sites, books, and dictionaries using different manual and machine learning techniques including web scraping and optical character recognition. OdiEnCorp 2.0 served in WAT 2020 EnglishOdia Indic Task.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-or_odiencorp.roots_fr_wikipediaROOTS Subset: roots_fr_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_fr_wikipedia.roots_pt_wikipediaROOTS Subset: roots_pt_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_pt_wikipedia.roots_ar_wikipediaROOTS Subset: roots_ar_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_wikipedia.roots_ca_wikipediaROOTS Subset: roots_ca_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_wikipedia.LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks
Dataset Card for "LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks"
More Information needed
roots_ca_enriched_conllu_ancora_for_ml_trainingROOTS Subset: roots_ca_enriched_conllu_ancora_for_ml_training
Enriched CONLLU Ancora for ML training
Dataset uid: enriched_conllu_ancora_for_ml_training
Description
This is an enriched version for Machine Learning purposes of the CONLLU adaptation of AnCora corpus .
This version of the corpus was developed by BSC TeMU as part of the AINA project, and has been used to do multi-task learning for the Catalan language Spacy 3.0 models.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_enriched_conllu_ancora_for_ml_training.roots_en_wikinewsROOTS Subset: roots_en_wikinews
wikinews_filtered
Dataset uid: wikinews_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0307 % of total
0.0701 % of ar
0.3036 % of pt
0.0271 % of en
0.0405 % of fr
0.2119 % of indic-ta
0.0081 % of zh
0.0510 % of es
0.0725 % of ca
BigScience processing steps
Filters applied to: ar
filter_wiki_user_titles
filter_wiki_non_text_type
dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikinews.roots_indic-hi_wikipediaROOTS Subset: roots_indic-hi_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-hi_wikipedia.roots_ar_uncorpusROOTS Subset: roots_ar_uncorpus
uncorpus
Dataset uid: uncorpus
Description
Homepage
Licensing
Speaker Locations
Sizes
2.8023 % of total
10.7390 % of ar
5.7970 % of fr
9.7477 % of es
2.0417 % of en
1.2540 % of zh
BigScience processing steps
Filters applied to: ar
dedup_document
filter_remove_empty_docs
filter_small_docs_bytes_300
Filters applied to: fr
dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_uncorpus.roots_es_uncorpusROOTS Subset: roots_es_uncorpus
uncorpus
Dataset uid: uncorpus
Description
Homepage
Licensing
Speaker Locations
Sizes
2.8023 % of total
10.7390 % of ar
5.7970 % of fr
9.7477 % of es
2.0417 % of en
1.2540 % of zh
BigScience processing steps
Filters applied to: ar
dedup_document
filter_remove_empty_docs
filter_small_docs_bytes_300
Filters applied to: fr
dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_uncorpus.roots_en_the_pile_usptoROOTS Subset: roots_en_the_pile_uspto
the_pile_uspto
Dataset uid: the_pile_uspto
Description
Homepage
Licensing
Speaker Locations
Sizes
0.5358 % of total
2.9032 % of en
BigScience processing steps
Filters applied to: en
dedup_document
filter_remove_empty_docs
filter_small_docs_bytes_1024
roots_zh_wikivoyageROOTS Subset: roots_zh_wikivoyage
wikivoyage_filtered
Dataset uid: wikivoyage_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0334 % of total
0.1097 % of en
0.0432 % of fr
0.0863 % of es
0.0084 % of zh
0.0892 % of vi
0.0464 % of indic-bn
0.0443 % of pt
0.0130 % of indic-hi
BigScience processing steps
Filters applied to: en
filter_wiki_user_titles
filter_wiki_non_text_type… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh_wikivoyage.roots_ar_openiti_procROOTS Subset: roots_ar_openiti_proc
OpenITI
Dataset uid: openiti_proc
Description
A corpus of Arabic texts that collected from Islamic books from different websites.
Homepage
https://zenodo.org/record/4075046
Licensing
non-commercial use
cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International
By exercising the Licensed Rights (defined below), You accept and agree to be bound by the terms and conditions of this… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_openiti_proc.roots_en_wikibooksROOTS Subset: roots_en_wikibooks
wikibooks_filtered
Dataset uid: wikibooks_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0897 % of total
0.2591 % of en
0.0965 % of fr
0.1691 % of es
0.2834 % of indic-hi
0.2172 % of pt
0.0149 % of zh
0.0279 % of ar
0.1374 % of vi
0.5025 % of id
0.3694 % of indic-ur
0.5744 % of eu
0.0769 % of ca
0.0519 % of indic-ta
0.1470 % of indic-mr
0.0751 % of indic-te
0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikibooks.roots_zh_wikibooksROOTS Subset: roots_zh_wikibooks
wikibooks_filtered
Dataset uid: wikibooks_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0897 % of total
0.2591 % of en
0.0965 % of fr
0.1691 % of es
0.2834 % of indic-hi
0.2172 % of pt
0.0149 % of zh
0.0279 % of ar
0.1374 % of vi
0.5025 % of id
0.3694 % of indic-ur
0.5744 % of eu
0.0769 % of ca
0.0519 % of indic-ta
0.1470 % of indic-mr
0.0751 % of indic-te
0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh_wikibooks.roots_zh-tw_wikipediaROOTS Subset: roots_zh-tw_wikipedia
wikipedia
Dataset uid: wikipedia
Description
Homepage
Licensing
Speaker Locations
Sizes
3.2299 % of total
4.2071 % of en
5.6773 % of ar
3.3416 % of fr
5.2815 % of es
12.4852 % of ca
0.4288 % of zh
0.4286 % of zh
5.4743 % of indic-bn
8.9062 % of indic-ta
21.3313 % of indic-te
4.4845 % of pt
4.0493 % of indic-hi
11.3163 % of indic-ml
22.5300 % of indic-ur
4.4902 % of vi
16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-tw_wikipedia.roots_en_ted_talks_iwsltROOTS Subset: roots_en_ted_talks_iwslt
WIT Ted Talks
Dataset uid: ted_talks_iwslt
Description
The Web Inventory Talk is a collection of the original Ted talks and their translated version. The translations are available in more than 109+ languages, though the distribution is not uniform.
Homepage
https://github.com/huggingface/datasets/blob/master/datasets/ted_talks_iwslt/README.md
Licensing
open license
cc-by-nc-4.0: Creative Commons Attribution… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_ted_talks_iwslt.roots_vi_vietnamese_poetryROOTS Subset: roots_vi_vietnamese_poetry
Vietnamese poetry from fsoft AI lab
Dataset uid: vietnamese_poetry
Description
171188 poems with different genres: luc-bat, 5-chu, 7-chu, 8-chu, 4-chu
Homepage
https://github.com/fsoft-ailab/Poem-Generator#dataset
Licensing
open license
mit: MIT License
Speaker Locations
South-eastern Asia
Vietnam
Sizes
0.0127 % of total
0.9285 % of vi
BigScience processing steps… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_vi_vietnamese_poetry.
