CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes230k downloads3y agoHugging Face02ljnlonoljpiljm /BIGstockimage-1.5Mimage1M<n<10M0 likes687 downloads1y agoHugging Face03ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-twoimage100K<n<1M0 likes493 downloads1y agoHugging Face04ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-oneimage100K<n<1M1 likes422 downloads1y agoHugging Face05bigstupidhats /openai_MMMLU_zhotext10K<n<100K0 likes405 downloads2y agoHugging Face06Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes404 downloads4y agoHugging Face07bigshishiga /dclm-baseline-subsettext1M<n<10M1 likes301 downloads1y agoHugging Face08bigstupidhats /openai_MMMLU_engtextn<1K0 likes273 downloads2y agoHugging Face09bigstupidhats /openai_MMMLU_arbtext10K<n<100K0 likes185 downloads2y agoHugging Face10bigstupidhats /openai_MMMLU_hintext10K<n<100K0 likes155 downloads2y agoHugging Face11bigstupidhats /dynasample_traintabular100K<n<1M0 likes144 downloads2y agoHugging Face12HF-SSSVVVTTT /bigsurvey_with_sent_srl_scorestext1K<n<10K0 likes114 downloads8d agoHugging Face13bigstupidhats /dynasample_multitasks_cleantabular1M<n<10M0 likes113 downloads2y agoHugging Face14bigstupidhats /openai_MMMLU_spatext10K<n<100K0 likes89 downloads2y agoHugging Face15bigstupidhats /dynasample_train_scoreby3llmstabular100K<n<1M0 likes85 downloads2y agoHugging Face16bigstupidhats /openai_MMMLU_rustext10K<n<100K0 likes83 downloads2y agoHugging Face17bigstupidhats /openai_MMMLU_swatext10K<n<100K0 likes82 downloads2y agoHugging Face18bigscience-data /roots_en_wikipediagatedROOTS Subset: roots_en_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.text1M<n<10M5 likes67 downloads4y agoHugging Face19bigscience-historical-texts /Open_Medieval_French Open Medieval French Source: https://github.com/OpenMedFr/texts text1K<n<10K3 likes60 downloads4y agoHugging Face20bigscience-data /roots_en_no_code_stackexchangegatedROOTS Subset: roots_en_no_code_stackexchange Stack Exchange Website Dataset uid: no_code_stackexchange Description Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers. Homepage https://stackexchange.com/ Licensing open license cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.text1M<n<10M1 likes56 downloads4y agoHugging Face21HF-SSSVVVTTT /bigsurvey_with_srltext1K<n<10K0 likes36 downloads5mo agoHugging Face22bigstupidhats /openai_MMMLU_deutext10K<n<100K0 likes33 downloads2y agoHugging Face23bigscience-data /roots_indic-bn_wikipediagatedROOTS Subset: roots_indic-bn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-bn_wikipedia.text100K<n<1M2 likes32 downloads4y agoHugging Face24bigscience-data /roots_zh-cn_wikipediagatedROOTS Subset: roots_zh-cn_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_zh-cn_wikipedia.text100K<n<1M32 likes30 downloads4y agoHugging Face25bigstupidhats /ultralinktext1M<n<10M0 likes29 downloads2y agoHugging Face26HF-SSSVVVTTT /bigsurvey_with_srl_newtext1K<n<10K0 likes28 downloads5mo agoHugging Face27bigscience-data /roots_es_wikipediagatedROOTS Subset: roots_es_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_es_wikipedia.text100K<n<1M0 likes27 downloads4y agoHugging Face28bigscience-data /roots_en_book_dash_booksgatedROOTS Subset: roots_en_book_dash_books Book Dash Books Dataset uid: book_dash_books Description Book Dash believes that every child should own one hundred books by the age of five. To that end, we gather creative professionals who volunteer to create new, African storybooks that anyone can freely translate, print and distribute. In this way, we have vastly reduced the costs involved in putting high-quality books in children’s hands and hearts. Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_book_dash_books.textn<1K2 likes25 downloads4y agoHugging Face29Tutoruslabs /bi-gsm8k Bi-GSM8K A high-quality bilingual dataset of 500 elementary-level math problems with step-by-step correct solutions and annotated student error patterns in both English and Korean. 📋 Overview Bi-GSM8K is designed to support intelligent tutoring systems by providing: Teacher-authored correct solutions with step-level details Expert-created simulated student solutions reflecting common mathematical misconceptions actually observed in real students Explicit step-level… See the full description on the dataset page: https://huggingface.co/datasets/Tutoruslabs/bi-gsm8k.text1K<n<10K0 likes25 downloads8mo agoHugging Face30bigscience-data /roots_fr_wikipediagatedROOTS Subset: roots_fr_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_fr_wikipedia.text100K<n<1M1 likes22 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.