CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZaMinVo /MultiviewX_Labelstabular10K<n<100K0 likes4.5k downloads7h agoHugging Face02electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.8k downloads2mo agoHugging Face03Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes706 downloads1y agoHugging Face04MultiSense /SaleData SalesLLM-10k (SaleData) ⭐ If you find this project helpful, please give us a star on GitHub! It means a lot to us. The official data repository of SalesLLM: Benchmarking LLM Realistic Selling Skill — accepted by EMNLP 2026 as a Main Paper. This repository hosts the SalesLLM-10k dataset: 10,000 high-quality, multi-turn sales conversations in Chinese across financial services (bank deposits, insurance, fund investment, stocks) and consumer products. 🔗 Related… See the full description on the dataset page: https://huggingface.co/datasets/MultiSense/SaleData.tabular10K<n<100K0 likes474 downloads19d agoHugging Face05nvidia /Nemotron-RL-Instruction-Following-MultiTurnChat-v1 Dataset Description: The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tabular1K<n<10K4 likes388 downloads7mo agoHugging Face06inclusionAI /MultiEditgated 🧩 MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks 📃 Arxiv 🚀 Dataset Overview Based on our MLLM-driven data construction pipeline using GPT-4o and GPT-Image-1, we introduce MultiEdit, a comprehensive large-scale instruction-based image editing dataset comprising over 107K samples targeting 6 challenging image editing tasks covering 56 subcategory editing types (18 non-style-transfer and 38 style transfer). We also release… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/MultiEdit.tabular100K<n<1M20 likes300 downloads1y agoHugging Face07ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes286 downloads5mo agoHugging Face08risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M3 likes276 downloads18d agoHugging Face09Nidhogg-zh /Multi-Querier_Dialogue Multi-Querier Dialogue (MQDialog) Dataset 📑 Dataset Details Dataset Description This is the dataset of "Querier-Aware LLM: Generating Personalized Responses to the Same Query from Different Queriers". The Multi-Querier Dialogue (MQDialog) dataset is designed to facilitate research in querier-aware personalization. It contains dialogues with various queriers for each reponder. The dataset is derived from English and Chinese scripts of popular TV shows and… See the full description on the dataset page: https://huggingface.co/datasets/Nidhogg-zh/Multi-Querier_Dialogue.tabular10K<n<100K5 likes262 downloads1y agoHugging Face10risaleinur /risale-nur-multilingual Risale-i Nur Multilingual Corpus Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir. Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.tabulartranslation100K<n<1M2 likes255 downloads1mo agoHugging Face11huggingface-projects /sd-multiplayer-dataTo access an image use the following Bucket URL: https://d26smi9133w0oo.cloudfront.net/ example: https://d26smi9133w0oo.cloudfront.net/room-7/1670520485-CZk4C72xBr5wPfTpwDAnG6-7648_7008-a-chicken-breaking-through-a-mirrornnotn.webp Bucket URL/key SQLite https://huggingface.co/datasets/huggingface-projects/sd-multiplayer-data/blob/main/rooms_data.db sqlite> PRAGMA table_info(rooms_data); 0|id|INTEGER|1||1 1|room_id|TEXT|1||0 2|uuid|TEXT|1||0 3|x|INTEGER|1||0 4|y|INTEGER|1||0 5|prompt|TEXT|1||0… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-projects/sd-multiplayer-data.tabular100K<n<1M2 likes227 downloads4y agoHugging Face12eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes224 downloads1y agoHugging Face13overthelex /multi-legal-bench Multi-Legal-Bench Paper: Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries (v3) Identical legal tasks evaluated on native court decisions from national registries in France, the Netherlands, Poland, the Czech Republic and Lithuania, with Ukrainian cells in the companion UA-Legal-Bench (not included here). Labels come from registry metadata. The v3 paper is an audit of what those labels let a benchmark measure: in… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/multi-legal-bench.tabulartext-classification10K<n<100K0 likes206 downloads4h agoHugging Face14nyu-dice-lab /lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private.tabular100K<n<1M0 likes185 downloads2y agoHugging Face15obaydata /multi-image-composition-instruction-following Multi-Image Composition Instruction-Following A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image. Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.imageimage-to-imagen<1K0 likes176 downloads6mo agoHugging Face16PaDaS-Lab /nfqa-multilingual-dataset NFQA Multilingual Dataset A large-scale multilingual dataset for Non-Factoid Question Answering (NFQA) classification, covering 49 languages and 8 question categories. Dataset Statistics Split Examples Train 28,653 Validation 3,539 Test 3,671 Total (Balanced) 35,863 Full Dataset (High Quality) 63,647 Dataset Composition Languages (49 total) Arabic (ar), Azerbaijani (az), Bulgarian (bg), Bengali (bn), Catalan (ca)… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/nfqa-multilingual-dataset.tabulartext-classification10K<n<100K1 likes158 downloads6mo agoHugging Face17bingran-you /multiplexing-artiq-results Multiplexing ARTIQ experimental results This dataset is a public snapshot of the HDF5 result files from artiq_working_dir - 20240315/results, collected for trapped-ion multiplexing and quantum-networking experiments. Contents artiq-results.tar.zst: 7,433 HDF5 result files organized by date and hour. MANIFEST.sha256: SHA-256 digest for every extracted HDF5 file. DATASET_INFO.json: source inventory and publication audit summary. The data span 170 date directories… See the full description on the dataset page: https://huggingface.co/datasets/bingran-you/multiplexing-artiq-results.tabularn<1K0 likes143 downloads1mo agoHugging Face18June30916 /multimodality-poc-llama31-ruler16k Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K Raw pre-RoPE query and hidden-state tensors captured during prefill, used to study whether the per-(layer, kv_head) query distribution is unimodal Gaussian (the assumption underpinning Expected Attention's MGF closed-form in kvpress). What's in here 65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5 prompts). Each file (~414 MB) contains: field dtype shape meaning hidden float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.tabularfeature-extractionn<1K0 likes142 downloads5mo agoHugging Face19fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes141 downloads8d agoHugging Face20amsminn /smash-karts-multiplayer-trajectory Smash Karts 멀티플레이 구현해줘 A single Codex coding-agent session implementing a multiplayer browser-based 3D kart battle game inspired by Smash Karts. The request covers multiplayer play, game rules, weapons and effects, keyboard controls, research, and implementation. The trajectory records the development process, tool calls and results, validation work, and the final response. Field Value Session title Smash Karts 멀티플레이 구현해줘 Session ID… See the full description on the dataset page: https://huggingface.co/datasets/amsminn/smash-karts-multiplayer-trajectory.tabularn<1K0 likes136 downloads21d agoHugging Face21superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes128 downloads6mo agoHugging Face22fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes124 downloads3y agoHugging Face23nyu-dice-lab /lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28 Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private.tabular100K<n<1M0 likes102 downloads2y agoHugging Face24TaiMingLu /Multilingual-BenchmarkThese are the GSM8K and ARC dataset translated by Google Translate. BibTex @misc{lu2024languagecountslearnunlearn, title={Every Language Counts: Learn and Unlearn in Multilingual LLMs}, author={Taiming Lu and Philipp Koehn}, year={2024}, eprint={2406.13748}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2406.13748}, } tabularzero-shot-classification1M<n<10M3 likes98 downloads2y agoHugging Face25multimodalart /agent-spaces-tracestabularn<1K0 likes97 downloads5mo agoHugging Face26eQOURSE /multilingual-speech Multilingual Indian Conversational Speech A dataset of naturalistic, spontaneous two-speaker conversations across 13 Indian languages, with segment-level transcripts, speaker profiles, timestamps, and recording metadata. Designed for ASR, TTS, speaker diarization, and conversational speech research. Languages (13) Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punjabi, Tamil, Telugu, Urdu. Content Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.audioautomatic-speech-recognitionn<1K0 likes94 downloads3mo agoHugging Face27avewright /exp186-sf-multipv-2m exp186 Stockfish MultiPV soft targets (~2M) Local Stockfish 18 full-strength MultiPV harvest for chess transformer soft-policy training. Contents data/positions_*.jsonl — ~2,000,000 positions (400 shards) soft_cache.pt — training-ready tensor cache (compact move vocab 1968) Labeling Engine: Stockfish 18 (full strength, no UCI_LimitStrength) MultiPV: 8 Depths: sampled in [2, 8] (weighted toward mid depths) Soft probs: softmax(cp / τ) with τ=120… See the full description on the dataset page: https://huggingface.co/datasets/avewright/exp186-sf-multipv-2m.tabularother1M<n<10M0 likes81 downloads3mo agoHugging Face28empero-ai /tasklist-grok-multilingual-100000x-unfiltered TaskGen Dataset Generated with taskgen by empero-org Run Parameters Parameter Value Model grok-4-1-fast-reasoning Temperature 0.9 Total Tasks 83052 Concurrency 30 workers API Base https://api.x.ai/v1 Generated 2026-04-07 14:31:14 Budget Cap $15.0000 Multilingual Yes (en, de, fr, es, nl, zh, ar, ru) Language Distribution Language Code Tasks Arabic ar 10446 German de 10397 Dutch nl 10353 Spanish es 10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.tabular100K<n<1M2 likes76 downloads6mo agoHugging Face29nyu-dice-lab /lm-eval-results-MTSAIR-multi_verse_model-private Dataset Card for Evaluation run of MTSAIR/multi_verse_model Dataset automatically created during the evaluation run of model MTSAIR/multi_verse_model The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MTSAIR-multi_verse_model-private.tabular100K<n<1M0 likes71 downloads2y agoHugging Face30beaugogh /openorca-multiplechoice-10kA 10k subset of OpenOrca dataset, focusing on multiple choice questions. Credit to Tian Xia. tabular10K<n<100K5 likes69 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.