CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrimeIntellect /Multi-SWE-RL-Verified Multi-SWE-RL-Verified Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end. Default dataset of the multiswe_v1 taskset. Changes vs upstream Starting from the 4,703-row re-upload: C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.tabulartext-generation1K<n<10K4 likes5.6k downloads3mo agoHugging Face02Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes706 downloads1y agoHugging Face03MultiSynt /MT-Reasoning MultiSynt MultiSynt is an open multilingual synthetic dataset. The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses. lang rows prompt_tokens reasoning_tokens response_tokens total_tokens deu_Latn 17_354_716 1_873_153_732 26_010_932_738 14_862_651_336 42_746_737_806 fra_Latn 17_354_716 1_802_885_115 25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.tabulartext-generation100M<n<1B0 likes519 downloads7mo agoHugging Face04yjlee36 /knowchat-multi-turn-dialogues KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.tabularquestion-answeringn<1K3 likes447 downloads1mo agoHugging Face05Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K0 likes425 downloads17d agoHugging Face06risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M3 likes276 downloads18d agoHugging Face07risaleinur /risale-nur-multilingual Risale-i Nur Multilingual Corpus Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir. Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.tabulartranslation100K<n<1M2 likes255 downloads1mo agoHugging Face08AlienKevin /SWE-ZERO-multilang-300-trajectories SWE-ZERO-multilang-300-trajectories 300 execution-free agentic rollouts from ricdomolm/mini-coder-1.7b across all 20 programming languages in nebius/SWE-rebench-V2 (5 PRs per language × 3 rollouts per PR), generated as part of marin-community/marin#4653. Companion to the Python-only run: AlienKevin/SWE-ZERO-1k-trajectories-32k — 1,000 rollouts on 100 Python PRs (10 repos × 10 PRs × 10 rollouts) AlienKevin/SWE-ZERO-1k-trajectories— the original 8k-context Python baseline Each row… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/SWE-ZERO-multilang-300-trajectories.tabulartext-generationn<1K0 likes229 downloads6mo agoHugging Face09eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes224 downloads1y agoHugging Face10taozi555 /novel-multilingual WebNovel Multilingual Dataset Dataset Description This dataset contains web novels scraped from WebNovel.com across multiple languages. Each entry includes the complete novel content with chapter information, metadata, and classification tags. Note: This dataset excludes content in the following languages: id, ID Dataset Statistics Total Novels: 8,324 Total Chapters: 233,410 Total Characters: 1,617,589,129 Languages: 10 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-multilingual.tabulartext-generation1K<n<10K2 likes198 downloads1y agoHugging Face11The-CoLab /multilingual-textarena-ColonelBlotto-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-ColonelBlotto-v0-train.tabulartext-generation1M<n<10M0 likes168 downloads2mo agoHugging Face12argilla /ultrafeedback-multi-binarized-preferences-cleaned UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned, and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the AlpacaEval, MT-Bench, and LM Eval Harness benchmarks. Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.tabulartext-generation100K<n<1M7 likes157 downloads3y agoHugging Face13The-CoLab /multilingual-textarena-SimpleTak-v0-train-v2 TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.tabulartext-generation100K<n<1M0 likes147 downloads3mo agoHugging Face14MauroPello /multilingual-reasoning-gym-sft Reasoning Gym SFT Dataset This dataset contains Supervised Fine-Tuning (SFT) reasoning data procedurally generated using Reasoning Gym environments. It is designed to train reasoning models (such as DeepSeek-R1-style or Qwen-Coder-style models) to explain their step-by-step reasoning chain before outputting a final answer wrapped inside LaTeX \boxed{...}. Where Does This Dataset Come From? This dataset is procedurally generated from Reasoning Gym, an open-source… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/multilingual-reasoning-gym-sft.tabulartext-generation100K<n<1M1 likes139 downloads3mo agoHugging Face15siddharthmb /multimodel-capitulation-interp Multi-model wrongful-capitulation internal-readout dataset Per-turn internal readouts + behavioral labels from two-model collaborative conversations (Qwen2.5-3B-Instruct × gemma-2-2b-it) on 6 reasoning benchmarks, restricted to the disagreement subset (one model right, one wrong solo). Built to test whether a linear correctness probe on the residual stream can predict wrongful capitulation (a model abandoning an answer it knew was correct under a partner's wrong assertion)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/multimodel-capitulation-interp.tabulartext-generation10K<n<100K0 likes138 downloads2mo agoHugging Face16BUAIR /Uganda-Multilingual-QA BUAIR Uganda Multilingual Q&A Parallel question–answer dataset for Ugandan languages, curated by the BUAIR Voice initiative at Busitema University. Dataset version: v2 (updated 2026-08-31) Each language config contains the same 4,256 agriculture / rural-livelihood Q&A pairs, with English as the shared source and translations into Japadhola, Ateso, Runyankore, and Luganda. Changelog (v2) Replaced v1 data (5,000 pairs from Multiligual-QA.xlsx) with cleaned data… See the full description on the dataset page: https://huggingface.co/datasets/BUAIR/Uganda-Multilingual-QA.tabularquestion-answering10K<n<100K0 likes131 downloads26d agoHugging Face17superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes128 downloads6mo agoHugging Face18The-CoLab /multilingual-textarena-ColonelBlotto-v0-train-v2 TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-ColonelBlotto-v0-train-v2.tabulartext-generation100K<n<1M0 likes120 downloads3mo agoHugging Face19The-CoLab /multilingual-textarena-KuhnPoker-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-KuhnPoker-v0-train.tabulartext-generation100K<n<1M0 likes118 downloads2mo agoHugging Face20gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes115 downloads2y agoHugging Face21The-CoLab /multilingual-textarena-SimpleTak-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.tabulartext-generation1M<n<10M0 likes107 downloads2mo agoHugging Face22eagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes98 downloads4mo agoHugging Face23projetogabi /csdi-multilingual CSDI: A Fine-Grained Fundus Image Dataset of Cataract Severity and Diagnostic Images Acknowledgment: This dataset is a multilingual extension of the original CSDI Cataract Diagnosis Dataset created by Xie, Z., Ao, M., Tang, H. et al. The original dataset provides expert written reports and diagnostic descriptions in English and Chinese. This version expands the diagnostic text into 31 additional languages to support cross-lingual research in automated cataract screening and… See the full description on the dataset page: https://huggingface.co/datasets/projetogabi/csdi-multilingual.tabularimage-text-to-text1K<n<10K0 likes96 downloads3mo agoHugging Face24beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes93 downloads1mo agoHugging Face25The-CoLab /multilingual-textarena-Nim-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset = load_dataset("The-CoLab/multilingual-textarena-Nim-v0-train"… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-Nim-v0-train.tabulartext-generation100K<n<1M0 likes92 downloads2mo agoHugging Face26neurips26 /MultimodalUnlearningEvalBenchmark 🧠 Multimodal Unlearning Evaluation Benchmark 📌 Overview This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning. It supports reproducibility of results in: Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026) 📊 Contents File Description 📄 multimodal_results.json Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench) 📄 unimodal_results.json CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.tabulartext-generationn<1K1 likes86 downloads5mo agoHugging Face27The-CoLab /multilingual-textarena-TicTacToe-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-TicTacToe-v0-train.tabulartext-generation100K<n<1M0 likes78 downloads2mo agoHugging Face28TnT /Multi_CodeNet4Repairtabulartext-generation100K<n<1M9 likes71 downloads3y agoHugging Face29The-CoLab /multilingual-textarena-Nim-v0-train-v2 TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-Nim-v0-train-v2.tabulartext-generation100K<n<1M0 likes67 downloads3mo agoHugging Face30wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes67 downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.