CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes27k downloads2y agoHugging Face02BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes7k downloads1y agoHugging Face03AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes6.3k downloads1y agoHugging Face04BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging Face05seedboxai /multitask_german_examples_32ktabular100K<n<1M15 likes560 downloads3y agoHugging Face06PreFLMR /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRtext10M<n<100M0 likes357 downloads2y agoHugging Face07Kowsher /multitask_vqa_benchmarkThis dataset is a part of . 🍈 MMT-47: Multimodal Multi-Task Benchmark 47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text) Cite our ICML-2026 paper for this dataset @article{kowsher2026lime, title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning}, author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_vqa_benchmark.image10K<n<100K0 likes345 downloads4mo agoHugging Face08WaltonFuture /VQA-MultiTaskimage100K<n<1M1 likes324 downloads1y agoHugging Face09mesolitica /Sampling-Multitask-National-Speech-Corpus-v1 Sampling Multitask-National-Speech-Corpus-v1 Original dataset from https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1, we only take Part 3 and do sampling. how to prepare the dataset huggingface-cli download \ mesolitica/Sampling-Multitask-National-Speech-Corpus-v1 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Sampling-Multitask-National-Speech-Corpus-v1.audio100K<n<1M0 likes121 downloads1y agoHugging Face10bigstupidhats /dynasample_multitasks_cleantabular1M<n<10M0 likes113 downloads2y agoHugging Face11vicgalle /configurable-system-prompt-multitask Configurable System Prompt Multi-task Dataset 🛞 We release the synthetic dataset for the multi-task experiments from the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data", https://huggingface.co/papers/2404.00495. This dataset has two sources for the examples: Self-critique on a safety task from Harmful Behaviours, using the SOLAR-Instruct model. It employs two system prompts to learn the different behaviors: You are a helpful yet harmless… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/configurable-system-prompt-multitask.texttext-generation1K<n<10K29 likes100 downloads2y agoHugging Face12Lelonthecodeur /multi-task-dataset Multi-Task Dataset Description A large-scale multi-task dataset designed for training and evaluating AI models across reasoning, mathematics, code, research, verification, data analysis, and general problem solving. Content 100,000,001 examples 20+ task families English + French Train / Validation / Test splits Structured reasoning and verification signals Multiple difficulty levels OOD and generalization-oriented examples Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset.tabular100M<n<1B2 likes78 downloads8d agoHugging Face13chronbmm /sanskrit-multitask-devanagaritext1M<n<10M0 likes77 downloads2y agoHugging Face14kaustubhg73 /multilingual-multitask-refusal Multilingual Multitask Refusal A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels. English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json. Rows 211,320 English seeds 1,761 Languages 15 Tasks 8 Product 1… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/multilingual-multitask-refusal.texttext-generation100K<n<1M0 likes70 downloads17d agoHugging Face15chronbmm /sanskrit-multitasktext1M<n<10M0 likes68 downloads2y agoHugging Face16wcarvalho /Multitask_Preplay_JaxMaze_models Multitask Preplay — jaxmaze model data Data for the paper "Multitask Preplay" (PNAS). Analysis code: https://github.com/wcarvalho/multitask_preplay (branch pnas). Splits: qlearning, usfa, dyna, preplay, her, bfs, dfs, greedy_euclidean, memory_based_euclidean. Each split is also available as a top-level parquet file. tabular10K<n<100K0 likes64 downloads3mo agoHugging Face17wcarvalho /MultitaskPreplay_jaxmaze_human_dftabular10K<n<100K0 likes63 downloads1y agoHugging Face18rntc /tmp-multitask-en-clinical Dataset Card for "tmp-multitask-en-clinical" More Information needed text100K<n<1M0 likes62 downloads1y agoHugging Face19LoserLi /MultiTasks-v2 MultiTasks-v2 MultiTasks-v2 is a collection of twelve multimodal benchmark subsets normalized into a shared image-question-answer format. Each subset provides a train split and a test split. Dataset Structure Each example contains: id: a unique sample identifier in the form {Dataset}_{split}_{index}. images: a list containing one image. Images are stored as JPEG bytes. problem: the prompt shown to the model. answer: the target answer. RefAdv is the only subset… See the full description on the dataset page: https://huggingface.co/datasets/LoserLi/MultiTasks-v2.image10K<n<100K1 likes59 downloads4mo agoHugging Face20wcarvalho /Multitask_Preplay_Craftax_models Multitask Preplay — craftax model data Data for the paper "Multitask Preplay" (PNAS). Analysis code: https://github.com/wcarvalho/multitask_preplay (branch pnas). Splits: qlearning, usfa, dyna, preplay, her, greedy_euclidean, memory_based_euclidean. Each split is also available as a top-level parquet file. tabular1K<n<10K0 likes54 downloads3mo agoHugging Face21Cache-SCA /UR7e_CaP_MultiTask_300epi_10fps UR7e CaP MultiTask 300epi 10fps This is a public LeRobot v3.0 multi-task dataset for UR7e Code-as-Policies manipulation. It merges three 100-episode 10fps datasets into one 300-episode training corpus while preserving the original numeric observations, actions, and task labels. The final uploaded copy stores all videos as 10fps H.264 MP4 files. Dataset Summary Robot: ur7e Format: LeRobot v3.0 FPS: 10 Episodes: 300 Frames: 152,926 Cameras:… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_CaP_MultiTask_300epi_10fps.tabularrobotics100K<n<1M0 likes52 downloads5mo agoHugging Face22riddickz /multitask_v1_toytextn<1K0 likes50 downloads1y agoHugging Face23fuyyy74 /SR-MultiTasktext10K<n<100K0 likes45 downloads3mo agoHugging Face24riddickz /multitask_v4_codingtabular10K<n<100K0 likes44 downloads1y agoHugging Face25Kowsher /multitask_textqa_benchmarkThis dataset is a part of . 🍈 MMT-47: Multimodal Multi-Task Benchmark 47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text) Cite our ICML-2026 paper for this dataset @article{kowsher2026lime, title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning}, author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_textqa_benchmark.text10K<n<100K0 likes44 downloads4mo agoHugging Face26Cache-SCA /UR7e_CaP_MultiTask_300epi_10fps_state_tplus1_action UR7e CaP MultiTask 300epi 10fps This is a public LeRobot v3.0 multi-task dataset for UR7e Code-as-Policies manipulation. It merges three 100-episode 10fps datasets into one 300-episode training corpus while preserving the original numeric observations, actions, and task labels. The final uploaded copy stores all videos as 10fps H.264 MP4 files. Dataset Summary Robot: ur7e Format: LeRobot v3.0 FPS: 10 Episodes: 300 Frames: 152,926 Cameras:… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_CaP_MultiTask_300epi_10fps_state_tplus1_action.tabularrobotics100K<n<1M0 likes42 downloads4mo agoHugging Face27riddickz /multitask_v3_coding_11ktabular10K<n<100K0 likes40 downloads1y agoHugging Face28nillohitroy /satquery-multitask-datasetimage10K<n<100K0 likes39 downloads19d agoHugging Face29AtesiT /ru-multitask-toxicity RU Multi-Task Toxicity Dataset Описание Датасет для обучения multi-task классификатора токсичности русскоязычных текстов по трём независимым бинарным категориям: profanity — ненормативная лексика (мат) threat — угрозы в адрес пользователя/третьих лиц illegal — запросы, связанные с нарушением закона (например, "как сделать бомбу", "где купить наркотики") Каждый текст может относиться сразу к нескольким категориям одновременно (multi-label задача), либо ни к одной… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-multitask-toxicity.tabulartext-classification10K<n<100K1 likes38 downloads2mo agoHugging Face30rntc /tmp-multitask-fr-clinical Dataset Card for "tmp-multitask-fr-clinical" More Information needed text100K<n<1M0 likes36 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.