CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes230k downloads3y agoHugging Face02ljnlonoljpiljm /BIGstockimage-1.5Mimage1M<n<10M0 likes687 downloads1y agoHugging Face03ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-twoimage100K<n<1M0 likes493 downloads1y agoHugging Face04ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-oneimage100K<n<1M1 likes422 downloads1y agoHugging Face05bigstupidhats /openai_MMMLU_zhotext10K<n<100K0 likes405 downloads2y agoHugging Face06Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes404 downloads4y agoHugging Face07Tristan /olm-october-2022-tokenized-1024-no-bigscience-filters Dataset Card for "olm-october-2022-tokenized-1024-no-bigscience-filters" More Information needed 10M<n<100M0 likes349 downloads4y agoHugging Face08bigshishiga /dclm-baseline-subsettext1M<n<10M1 likes301 downloads1y agoHugging Face09bigstupidhats /openai_MMMLU_engtextn<1K0 likes273 downloads2y agoHugging Face10bigstupidhats /openai_MMMLU_arbtext10K<n<100K0 likes185 downloads2y agoHugging Face11bigstupidhats /openai_MMMLU_hintext10K<n<100K0 likes155 downloads2y agoHugging Face12bigstupidhats /dynasample_traintabular100K<n<1M0 likes144 downloads2y agoHugging Face13HF-SSSVVVTTT /bigsurvey_with_sent_srl_scorestext1K<n<10K0 likes114 downloads7d agoHugging Face14bigstupidhats /dynasample_multitasks_cleantabular1M<n<10M0 likes113 downloads2y agoHugging Face15bigstupidhats /openai_MMMLU_spatext10K<n<100K0 likes89 downloads2y agoHugging Face16bigstupidhats /dynasample_train_scoreby3llmstabular100K<n<1M0 likes85 downloads2y agoHugging Face17bigstupidhats /openai_MMMLU_rustext10K<n<100K0 likes83 downloads2y agoHugging Face18bigstupidhats /openai_MMMLU_swatext10K<n<100K0 likes82 downloads2y agoHugging Face19bigscience-data /roots_en_wikipediagatedROOTS Subset: roots_en_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.text1M<n<10M5 likes67 downloads4y agoHugging Face20bigscience-historical-texts /Open_Medieval_French Open Medieval French Source: https://github.com/OpenMedFr/texts text1K<n<10K3 likes60 downloads4y agoHugging Face21bigscience-data /roots_en_no_code_stackexchangegatedROOTS Subset: roots_en_no_code_stackexchange Stack Exchange Website Dataset uid: no_code_stackexchange Description Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers. Homepage https://stackexchange.com/ Licensing open license cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.text1M<n<10M1 likes56 downloads4y agoHugging Face22yuan1119 /bigsmallglass_10ep_testtabular1K<n<10K0 likes50 downloads29d agoHugging Face23HF-SSSVVVTTT /bigsurvey_with_srltext1K<n<10K0 likes36 downloads5mo agoHugging Face24bigstupidhats /openai_MMMLU_deutext10K<n<100K0 likes33 downloads2y agoHugging Face25bigscience-data /roots_indic-bn_wikipediagatedROOTS Subset: roots_indic-bn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-bn_wikipedia.text100K<n<1M2 likes32 downloads4y agoHugging Face26bigscience-data /roots_zh-cn_wikipediagatedROOTS Subset: roots_zh-cn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-cn_wikipedia.text100K<n<1M32 likes30 downloads4y agoHugging Face27bigstupidhats /ultralinktext1M<n<10M0 likes29 downloads2y agoHugging Face28HF-SSSVVVTTT /bigsurvey_with_srl_newtext1K<n<10K0 likes28 downloads5mo agoHugging Face29bigscience-data /roots_es_wikipediagatedROOTS Subset: roots_es_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_wikipedia.text100K<n<1M0 likes27 downloads4y agoHugging Face30bigscience-data /roots_en_book_dash_booksgatedROOTS Subset: roots_en_book_dash_books Book Dash Books Dataset uid: book_dash_books Description Book Dash believes that every child should own one hundred books by the age of five. To that end, we gather creative professionals who volunteer to create new, African storybooks that anyone can freely translate, print and distribute. In this way, we have vastly reduced the costs involved in putting high-quality books in children’s hands and hearts. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_book_dash_books.textn<1K2 likes25 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.