CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes229k downloads3y agoHugging Face02Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes399 downloads4y agoHugging Face03Tristan /olm-october-2022-tokenized-1024-no-bigscience-filters Dataset Card for "olm-october-2022-tokenized-1024-no-bigscience-filters" More Information needed 10M<n<100M0 likes351 downloads4y agoHugging Face04bigscience-data /roots_en_wikipediagatedROOTS Subset: roots_en_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.text1M<n<10M5 likes68 downloads4y agoHugging Face05bigscience-historical-texts /Open_Medieval_French Open Medieval French Source: https://github.com/OpenMedFr/texts text1K<n<10K3 likes57 downloads4y agoHugging Face06bigscience-data /roots_en_no_code_stackexchangegatedROOTS Subset: roots_en_no_code_stackexchange Stack Exchange Website Dataset uid: no_code_stackexchange Description Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers. Homepage https://stackexchange.com/ Licensing open license cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.text1M<n<10M1 likes55 downloads4y agoHugging Face07bigscience-data /roots_indic-bn_wikipediagatedROOTS Subset: roots_indic-bn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-bn_wikipedia.text100K<n<1M2 likes32 downloads4y agoHugging Face08bigscience-data /roots_zh-cn_wikipediagatedROOTS Subset: roots_zh-cn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-cn_wikipedia.text100K<n<1M32 likes30 downloads4y agoHugging Face09bigscience-data /roots_es_wikipediagatedROOTS Subset: roots_es_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_wikipedia.text100K<n<1M0 likes27 downloads4y agoHugging Face10bigscience-data /roots_en_book_dash_booksgatedROOTS Subset: roots_en_book_dash_books Book Dash Books Dataset uid: book_dash_books Description Book Dash believes that every child should own one hundred books by the age of five. To that end, we gather creative professionals who volunteer to create new, African storybooks that anyone can freely translate, print and distribute. In this way, we have vastly reduced the costs involved in putting high-quality books in children’s hands and hearts. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_book_dash_books.textn<1K2 likes25 downloads4y agoHugging Face11bigscience-data /roots_en_odiencorpgatedROOTS Subset: roots_en_odiencorp OdiEnCorp2.0 Dataset uid: odiencorp Description OdiEnCorp is a collection of Odia-English parallel and Odia monolingual sentences collected from different sources such as Odia Wikipedia, web sites, books, and dictionaries using different manual and machine learning techniques including web scraping and optical character recognition. OdiEnCorp 2.0 served in WAT 2020 EnglishOdia Indic Task. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_odiencorp.textn<1K0 likes24 downloads4y agoHugging Face12bigscience-data /roots_indic-or_odiencorpgatedROOTS Subset: roots_indic-or_odiencorp OdiEnCorp2.0 Dataset uid: odiencorp Description OdiEnCorp is a collection of Odia-English parallel and Odia monolingual sentences collected from different sources such as Odia Wikipedia, web sites, books, and dictionaries using different manual and machine learning techniques including web scraping and optical character recognition. OdiEnCorp 2.0 served in WAT 2020 EnglishOdia Indic Task. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-or_odiencorp.text10K<n<100K0 likes23 downloads4y agoHugging Face13bigscience-data /roots_fr_wikipediagatedROOTS Subset: roots_fr_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_fr_wikipedia.text100K<n<1M1 likes22 downloads4y agoHugging Face14bigscience-data /roots_pt_wikipediagatedROOTS Subset: roots_pt_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_pt_wikipedia.text100K<n<1M2 likes20 downloads4y agoHugging Face15bigscience-data /roots_ar_wikipediagatedROOTS Subset: roots_ar_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_wikipedia.text1M<n<10M1 likes19 downloads4y agoHugging Face16bigscience-data /roots_ca_wikipediagatedROOTS Subset: roots_ca_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_wikipedia.text100K<n<1M0 likes19 downloads4y agoHugging Face17Multimodal-Fatima /LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks Dataset Card for "LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks" More Information needed text1K<n<10K0 likes17 downloads4y agoHugging Face18bigscience-data /roots_ca_enriched_conllu_ancora_for_ml_traininggatedROOTS Subset: roots_ca_enriched_conllu_ancora_for_ml_training Enriched CONLLU Ancora for ML training Dataset uid: enriched_conllu_ancora_for_ml_training Description This is an enriched version for Machine Learning purposes of the CONLLU adaptation of AnCora corpus . This version of the corpus was developed by BSC TeMU as part of the AINA project, and has been used to do multi-task learning for the Catalan language Spacy 3.0 models. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_enriched_conllu_ancora_for_ml_training.textn<1K0 likes16 downloads4y agoHugging Face19bigscience-data /roots_en_wikinewsgatedROOTS Subset: roots_en_wikinews wikinews_filtered Dataset uid: wikinews_filtered Description Homepage Licensing Speaker Locations Sizes 0.0307 % of total 0.0701 % of ar 0.3036 % of pt 0.0271 % of en 0.0405 % of fr 0.2119 % of indic-ta 0.0081 % of zh 0.0510 % of es 0.0725 % of ca BigScience processing steps Filters applied to: ar filter_wiki_user_titles filter_wiki_non_text_type dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikinews.text10K<n<100K0 likes16 downloads4y agoHugging Face20bigscience-data /roots_indic-hi_wikipediagatedROOTS Subset: roots_indic-hi_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-hi_wikipedia.text100K<n<1M0 likes16 downloads4y agoHugging Face21bigscience-data /roots_ar_uncorpusgatedROOTS Subset: roots_ar_uncorpus uncorpus Dataset uid: uncorpus Description Homepage Licensing Speaker Locations Sizes 2.8023 % of total 10.7390 % of ar 5.7970 % of fr 9.7477 % of es 2.0417 % of en 1.2540 % of zh BigScience processing steps Filters applied to: ar dedup_document filter_remove_empty_docs filter_small_docs_bytes_300 Filters applied to: fr dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_uncorpus.text100K<n<1M0 likes15 downloads4y agoHugging Face22bigscience-data /roots_es_uncorpusgatedROOTS Subset: roots_es_uncorpus uncorpus Dataset uid: uncorpus Description Homepage Licensing Speaker Locations Sizes 2.8023 % of total 10.7390 % of ar 5.7970 % of fr 9.7477 % of es 2.0417 % of en 1.2540 % of zh BigScience processing steps Filters applied to: ar dedup_document filter_remove_empty_docs filter_small_docs_bytes_300 Filters applied to: fr dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_uncorpus.text100K<n<1M1 likes14 downloads4y agoHugging Face23bigscience-data /roots_en_the_pile_usptogatedROOTS Subset: roots_en_the_pile_uspto the_pile_uspto Dataset uid: the_pile_uspto Description Homepage Licensing Speaker Locations Sizes 0.5358 % of total 2.9032 % of en BigScience processing steps Filters applied to: en dedup_document filter_remove_empty_docs filter_small_docs_bytes_1024 text1M<n<10M1 likes14 downloads4y agoHugging Face24bigscience-data /roots_zh_wikivoyagegatedROOTS Subset: roots_zh_wikivoyage wikivoyage_filtered Dataset uid: wikivoyage_filtered Description Homepage Licensing Speaker Locations Sizes 0.0334 % of total 0.1097 % of en 0.0432 % of fr 0.0863 % of es 0.0084 % of zh 0.0892 % of vi 0.0464 % of indic-bn 0.0443 % of pt 0.0130 % of indic-hi BigScience processing steps Filters applied to: en filter_wiki_user_titles filter_wiki_non_text_type… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh_wikivoyage.text1K<n<10K1 likes14 downloads4y agoHugging Face25bigscience-data /roots_ar_openiti_procgatedROOTS Subset: roots_ar_openiti_proc OpenITI Dataset uid: openiti_proc Description A corpus of Arabic texts that collected from Islamic books from different websites. Homepage https://zenodo.org/record/4075046 Licensing non-commercial use cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International By exercising the Licensed Rights (defined below), You accept and agree to be bound by the terms and conditions of this… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_openiti_proc.text1K<n<10K2 likes13 downloads4y agoHugging Face26bigscience-data /roots_en_wikibooksgatedROOTS Subset: roots_en_wikibooks wikibooks_filtered Dataset uid: wikibooks_filtered Description Homepage Licensing Speaker Locations Sizes 0.0897 % of total 0.2591 % of en 0.0965 % of fr 0.1691 % of es 0.2834 % of indic-hi 0.2172 % of pt 0.0149 % of zh 0.0279 % of ar 0.1374 % of vi 0.5025 % of id 0.3694 % of indic-ur 0.5744 % of eu 0.0769 % of ca 0.0519 % of indic-ta 0.1470 % of indic-mr 0.0751 % of indic-te 0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikibooks.text10K<n<100K0 likes13 downloads4y agoHugging Face27bigscience-data /roots_zh_wikibooksgatedROOTS Subset: roots_zh_wikibooks wikibooks_filtered Dataset uid: wikibooks_filtered Description Homepage Licensing Speaker Locations Sizes 0.0897 % of total 0.2591 % of en 0.0965 % of fr 0.1691 % of es 0.2834 % of indic-hi 0.2172 % of pt 0.0149 % of zh 0.0279 % of ar 0.1374 % of vi 0.5025 % of id 0.3694 % of indic-ur 0.5744 % of eu 0.0769 % of ca 0.0519 % of indic-ta 0.1470 % of indic-mr 0.0751 % of indic-te 0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh_wikibooks.text1K<n<10K9 likes13 downloads4y agoHugging Face28bigscience-data /roots_zh-tw_wikipediagatedROOTS Subset: roots_zh-tw_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-tw_wikipedia.text100K<n<1M14 likes13 downloads4y agoHugging Face29bigscience-data /roots_en_ted_talks_iwsltgatedROOTS Subset: roots_en_ted_talks_iwslt WIT Ted Talks Dataset uid: ted_talks_iwslt Description The Web Inventory Talk is a collection of the original Ted talks and their translated version. The translations are available in more than 109+ languages, though the distribution is not uniform. Homepage https://github.com/huggingface/datasets/blob/master/datasets/ted_talks_iwslt/README.md Licensing open license cc-by-nc-4.0: Creative Commons Attribution… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_ted_talks_iwslt.text1K<n<10K1 likes12 downloads4y agoHugging Face30bigscience-data /roots_vi_vietnamese_poetrygatedROOTS Subset: roots_vi_vietnamese_poetry Vietnamese poetry from fsoft AI lab Dataset uid: vietnamese_poetry Description 171188 poems with different genres: luc-bat, 5-chu, 7-chu, 8-chu, 4-chu Homepage https://github.com/fsoft-ailab/Poem-Generator#dataset Licensing open license mit: MIT License Speaker Locations South-eastern Asia Vietnam Sizes 0.0127 % of total 0.9285 % of vi BigScience processing steps… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_vi_vietnamese_poetry.text100K<n<1M6 likes12 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.