datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
P3
Dataset Card for P3
Dataset Summary
P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.xP3allxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.xP3mtxP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.BIGstockimage-1.5Mshades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.bigscience-lama
Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models.
@inproceedings{petroni2020how,
title={How Context Affects Language Models' Factual Predictions},
author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel},
booktitle={Automated Knowledge Base Construction},
year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.BIGstockimage-1.5M-scored-pt-twoBIGstockimage-1.5M-scored-pt-oneopenai_MMMLU_zhoolm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters"
More Information needed
big_setlmsys-chat-enbigslide
Dataset Card for Bigslide.ru Presentations
Dataset Summary
This dataset contains metadata and original files for 50,872 presentations from the bigslide.ru platform, a presentation storage and viewing service for school students. The dataset includes information such as presentation titles, URLs, download URLs, and extracted text content where available.
Languages
The dataset is multilingual, with Russian being the primary language. Other languages present… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/bigslide.dclm-baseline-subsetopenai_MMMLU_engopenai_MMMLU_arbwildchat-enopenai_MMMLU_hindynasample_trainlm_code_github-eval_subsetcollaborative_catalogbigsurvey_with_sent_srl_scoresdynasample_multitasks_cleanUltraMedicalMetaMathQAopenai_MMMLU_spadynasample_train_scoreby3llmsopenai_MMMLU_rusopenai_MMMLU_swafinancial-instruction-aq22
