CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes230k downloads3y agoHugging Face02bigscience /xP3allxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.textother10M<n<100M32 likes42k downloads3y agoHugging Face03bigscience /xP3mtxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.textother10M<n<100M26 likes19k downloads3y agoHugging Face04ljnlonoljpiljm /BIGstockimage-1.5Mimage1M<n<10M0 likes687 downloads1y agoHugging Face05bigscience-catalogue-data /shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades Data Statement for SHADES How to use this document: Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.text10K<n<100K4 likes682 downloads2y agoHugging Face06janck /bigscience-lama Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models. @inproceedings{petroni2020how, title={How Context Affects Language Models' Factual Predictions}, author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel}, booktitle={Automated Knowledge Base Construction}, year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.texttext-retrieval10K<n<100K1 likes497 downloads4y agoHugging Face07ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-twoimage100K<n<1M0 likes493 downloads1y agoHugging Face08ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-oneimage100K<n<1M1 likes422 downloads1y agoHugging Face09bigstupidhats /openai_MMMLU_zhotext10K<n<100K0 likes405 downloads2y agoHugging Face10Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes404 downloads4y agoHugging Face11G-reen /big_settext100K<n<1M0 likes390 downloads9mo agoHugging Face12bigstupidhats /lmsys-chat-entext100K<n<1M0 likes318 downloads2y agoHugging Face13nyuuzyou /bigslide Dataset Card for Bigslide.ru Presentations Dataset Summary This dataset contains metadata and original files for 50,872 presentations from the bigslide.ru platform, a presentation storage and viewing service for school students. The dataset includes information such as presentation titles, URLs, download URLs, and extracted text content where available. Languages The dataset is multilingual, with Russian being the primary language. Other languages present… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/bigslide.texttext-classification10K<n<100K0 likes303 downloads2y agoHugging Face14bigshishiga /dclm-baseline-subsettext1M<n<10M1 likes301 downloads1y agoHugging Face15bigstupidhats /openai_MMMLU_engtextn<1K0 likes273 downloads2y agoHugging Face16bigstupidhats /openai_MMMLU_arbtext10K<n<100K0 likes185 downloads2y agoHugging Face17bigstupidhats /wildchat-entabular100K<n<1M0 likes162 downloads2y agoHugging Face18bigstupidhats /openai_MMMLU_hintext10K<n<100K0 likes155 downloads2y agoHugging Face19bigstupidhats /dynasample_traintabular100K<n<1M0 likes144 downloads2y agoHugging Face20bigscience-catalogue-data-dev /lm_code_github-eval_subsettext10K<n<100K2 likes139 downloads5y agoHugging Face21bigscience /collaborative_catalogtextn<1K1 likes120 downloads4y agoHugging Face22HF-SSSVVVTTT /bigsurvey_with_sent_srl_scorestext1K<n<10K0 likes114 downloads7d agoHugging Face23bigstupidhats /dynasample_multitasks_cleantabular1M<n<10M0 likes113 downloads2y agoHugging Face24bigstupidhats /UltraMedicaltext100K<n<1M0 likes99 downloads2y agoHugging Face25bigstupidhats /MetaMathQAtext100K<n<1M0 likes90 downloads2y agoHugging Face26bigstupidhats /openai_MMMLU_spatext10K<n<100K0 likes89 downloads2y agoHugging Face27bigstupidhats /dynasample_train_scoreby3llmstabular100K<n<1M0 likes85 downloads2y agoHugging Face28bigstupidhats /openai_MMMLU_rustext10K<n<100K0 likes83 downloads2y agoHugging Face29bigstupidhats /openai_MMMLU_swatext10K<n<100K0 likes82 downloads2y agoHugging Face30bigstupidhats /financial-instruction-aq22text100K<n<1M0 likes77 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.