CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /evaluation-results@misc{muennighoff2022crosslingual, title={Crosslingual Generalization through Multitask Finetuning}, author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel}, year={2022}, eprint={2211.01786}, archivePrefix={arXiv}, primaryClass={cs.CL} }other100M<n<1B10 likes301k downloads3y agoHugging Face02bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes229k downloads3y agoHugging Face03bigscience /xP3xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.other100M<n<1B113 likes71k downloads3y agoHugging Face04bigscience /xP3allxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.textother10M<n<100M32 likes42k downloads3y agoHugging Face05bigscience /xP3mtxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.textother10M<n<100M26 likes19k downloads3y agoHugging Face06BigScienceBiasEval /crows_pairs_multilingualThis is a revised version of CrowS-Pairs that measures stereotypes in language modelling in both English and French.3 likes1.9k downloads3y agoHugging Face07bigscience /xP3megds Dataset Card for xP3 Dataset Summary xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot. Creation: The dataset can be recreated using instructions available here. We provide this version to save processing time and ease reproducibility. Languages: 46 (Can… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/xP3megds.other100M<n<1B3 likes1.4k downloads3y agoHugging Face08open-llm-leaderboard-old /details_bigscience__bloom-7b1 Dataset Card for Evaluation run of bigscience/bloom-7b1 Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloom-7b1 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-7b1.0 likes620 downloads3y agoHugging Face09applied-ai-018 /peacock-data-public-datasets-idc-bigscience bigscience Research workshop on large language models - The Summer of Language Models 21 At the moment we have 2 code repos: https://github.com/bigscience-workshop/Megatron-DeepSpeed - this is our flagship code base https://github.com/bigscience-workshop/bigscience - (this repo) for everything else - docs, experiments, etc. Currently, the most active segments of this repo are: JZ - Lots of information about our work environment which helps evaluate, plan and get things done… See the full description on the dataset page: https://huggingface.co/datasets/applied-ai-018/peacock-data-public-datasets-idc-bigscience.0 likes567 downloads2y agoHugging Face10bigscience-catalogue-data /shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades Data Statement for SHADES How to use this document: Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.text10K<n<100K4 likes525 downloads2y agoHugging Face11janck /bigscience-lama Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models. @inproceedings{petroni2020how, title={How Context Affects Language Models' Factual Predictions}, author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel}, booktitle={Automated Knowledge Base Construction}, year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.texttext-retrieval10K<n<100K1 likes463 downloads4y agoHugging Face12Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes399 downloads4y agoHugging Face13Tristan /olm-october-2022-tokenized-1024-no-bigscience-filters Dataset Card for "olm-october-2022-tokenized-1024-no-bigscience-filters" More Information needed 10M<n<100M0 likes351 downloads4y agoHugging Face14open-llm-leaderboard-old /details_bigscience__bloom Dataset Card for Evaluation run of None Dataset Summary Dataset automatically created during the evaluation run of model None on the Open LLM Leaderboard. The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom.0 likes187 downloads3y agoHugging Face15bigscience-catalogue-data-dev /lm_code_github-eval_subsettext10K<n<100K2 likes141 downloads5y agoHugging Face16bigscience /collaborative_catalogtextn<1K1 likes120 downloads4y agoHugging Face17open-llm-leaderboard-old /details_bigscience__bloom-560m Dataset Card for Evaluation run of bigscience/bloom-560m Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloom-560m on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 13 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-560m.0 likes103 downloads3y agoHugging Face18bigscience-historical-texts /HIPE2020_sent-splitTODO0 likes97 downloads4y agoHugging Face19bigscience-historical-texts /hipe2020TODO3 likes82 downloads4y agoHugging Face20open-llm-leaderboard-old /details_bigscience__bloom-1b1 Dataset Card for Evaluation run of bigscience/bloom-1b1 Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloom-1b1 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-1b1.0 likes73 downloads3y agoHugging Face21bigscience-data /roots_en_wikipediagatedROOTS Subset: roots_en_wikipedia wikipedia Dataset uid: wikipedia Description Homepage Licensing Speaker Locations Sizes 3.2299 % of total 4.2071 % of en 5.6773 % of ar 3.3416 % of fr 5.2815 % of es 12.4852 % of ca 0.4288 % of zh 0.4286 % of zh 5.4743 % of indic-bn 8.9062 % of indic-ta 21.3313 % of indic-te 4.4845 % of pt 4.0493 % of indic-hi 11.3163 % of indic-ml 22.5300 % of indic-ur 4.4902 % of vi 16.9916 % of indic-kn… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_wikipedia.text1M<n<10M5 likes68 downloads4y agoHugging Face22bigscience-historical-texts /Open_Medieval_French Open Medieval French Source: https://github.com/OpenMedFr/texts text1K<n<10K3 likes57 downloads4y agoHugging Face23open-llm-leaderboard-old /details_bigscience__bloomz-7b1 Dataset Card for Evaluation run of bigscience/bloomz-7b1 Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloomz-7b1 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloomz-7b1.0 likes57 downloads3y agoHugging Face24bigscience-data /roots_en_no_code_stackexchangegatedROOTS Subset: roots_en_no_code_stackexchange Stack Exchange Website Dataset uid: no_code_stackexchange Description Launched in 2010, the Stack Exchange network comprises 173 Q&A communities including Stack Overflow, the largest, most trusted online community for developers to learn, share their knowledge, and build their careers. Homepage https://stackexchange.com/ Licensing open license cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_en_no_code_stackexchange.text1M<n<10M1 likes55 downloads4y agoHugging Face25bigscience /tokenizer-probing-ud2.101 likes52 downloads4y agoHugging Face26open-llm-leaderboard /bigscience__bloom-7b1-detailsgated Dataset Card for Evaluation run of bigscience/bloom-7b1 Dataset automatically created during the evaluation run of model bigscience/bloom-7b1 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-7b1-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face27open-llm-leaderboard-old /details_bigscience__bloom-1b7 Dataset Card for Evaluation run of bigscience/bloom-1b7 Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloom-1b7 on the Open LLM Leaderboard. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloom-1b7.0 likes44 downloads3y agoHugging Face28open-llm-leaderboard-old /details_bigscience__bloomz-3b Dataset Card for Evaluation run of bigscience/bloomz-3b Dataset Summary Dataset automatically created during the evaluation run of model bigscience/bloomz-3b on the Open LLM Leaderboard. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigscience__bloomz-3b.0 likes43 downloads3y agoHugging Face29open-llm-leaderboard /bigscience__bloom-3b-detailsgated Dataset Card for Evaluation run of bigscience/bloom-3b Dataset automatically created during the evaluation run of model bigscience/bloom-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-3b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face30open-llm-leaderboard /bigscience__bloom-560m-detailsgated Dataset Card for Evaluation run of bigscience/bloom-560m Dataset automatically created during the evaluation run of model bigscience/bloom-560m The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-560m-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.