CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChilleD /MultiArithtextn<1K17 likes103k downloads3y agoHugging Face02Mutonix /Vript_Multilingual 🎬 Vript: A Video Is Worth Thousands of Words [Github Repo] We construct another fine-grained video-text dataset with 19.1K annotated high-resolution UGC videos (~677k clips) in multiple languages to be the Vript_Multilingual. New in Vript_Multilingual: Multilingual: zh (60%), en (17%), de (15%), ja (6%), ko (2%), ru (<1%), es (<1%), pt (<1%), jv (<1%), fr (<1%), id (<1%), vi (<1%) More diverse and fine-grained categories: 113 categories (please check vript_CN-V2_meta.json)… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Multilingual.textvideo-classification100K<n<1M7 likes14k downloads2y agoHugging Face03alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes6.8k downloads2mo agoHugging Face04yixuantt /MultiHopRAG Dataset Card for Dataset Name A Dataset for Evaluating Retrieval-Augmented Generation Across Documents Dataset Description MultiHop-RAG: a QA dataset to evaluate retrieval and reasoning across documents with metadata in the RAG pipelines. It contains 2556 queries, with evidence for each query distributed across 2 to 4 documents. The queries also involve document metadata, reflecting complex scenarios commonly found in real-world RAG applications. Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/yixuantt/MultiHopRAG.textquestion-answering1K<n<10K74 likes6.2k downloads3y agoHugging Face05zwcolin /clevr-multichange CLEVR-Multi-Change (30–40 objects) Two-image change-captioning data used in "Stateful Visual Encoders for Vision-Language Models" (the Multi-object Visual Differencing task). Each example is a before/after pair of a CLEVR scene with 30–40 objects and 4 simultaneous changes (add / delete / move / replace), rendered at 768×768 with a wide camera angle. Built with the CLEVR-Multi-Change engine (Johnson et al. 2017; Qiu et al. 2021). Code & paper:… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/clevr-multichange.imageimage-to-text100K<n<1M0 likes6.1k downloads4mo agoHugging Face06ZaMinVo /MultiviewX_Labelstabular10K<n<100K0 likes4.5k downloads2h agoHugging Face07Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes4.2k downloads2y agoHugging Face08Daoguang /Multi-SWE-bench SWE-bench-Java: A GitHub Issue Resolving Benchmark for Java 📰 News [Aug. 27, 2024]:We’ve released the JAVA version of SWE-bench! Check it out on Hugging Face. For more details, see our paper! 📄 Abstract GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/Daoguang/Multi-SWE-bench.textn<1K7 likes4.1k downloads2y agoHugging Face09Multilingual-Multimodal-NLP /IfEvalCode-testsettextn<1K2 likes3.6k downloads1y agoHugging Face10SetFit /amazon_reviews_multi_entext100K<n<1M7 likes3k downloads4y agoHugging Face11allenai /multilingual_mbppMBPP translated to 15 programming languages using o4-mini-medium. source_language = "python" target_languages = [ "cpp", "c", "javascript", "java", "php", "csharp", "typescript", "bash", "swift", "go", "rust", "ruby", "r", "matlab", "scala", "haskell" ] effort = "medium" dataset_name = "google-research-datasets/mbpp" model = "o4-mini" text10K<n<100K2 likes3k downloads1y agoHugging Face12electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.8k downloads2mo agoHugging Face13nvidia /Nemotron-SFT-Multilingual-v2 Dataset Description: Nemotron-SFT-Multilingual-v2 is a multilingual supervised fine-tuning (SFT) dataset for post-training text-generation models. It is generated by translating seed data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1, adding multilingual coverage for Hindi (hi), Korean (ko), Brazilian Portuguese (pt-br), and refreshed Japanese (ja) data. The dataset is generated with a new data processing pipeline that avoids line-breaking… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v2.texttext-generation100K<n<1M14 likes2.5k downloads4mo agoHugging Face14bentrevett /multi30k Multi30k This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here. Each example consists of an "en" and a "de" feature. "en" is an English sentence, and "de" is the German translation of the English sentence. Data Splits The Multi30k dataset has 3 splits: train, validation, and test. Dataset Split Number of Instances in Split Train 29,000 Validation 1,014 Test 1,000 Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/bentrevett/multi30k.texttranslation10K<n<100K19 likes2.1k downloads4y agoHugging Face15Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2k downloads8mo agoHugging Face16LianeMarilin /CADBench-Extended-Multimodal-Dataset Dataset Card Dataset Description CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics. Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.3dimage-to-textn<1K2 likes2k downloads24d agoHugging Face17ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face18Babelscape /multinerd Dataset Card for MultiNERD dataset Description Summary: In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/multinerd.texttoken-classification1M<n<10M26 likes1.6k downloads3y agoHugging Face19jinggu /MultipanelVQAimagen<1K0 likes1.3k downloads3y agoHugging Face20dlwh /MultiLegalPile_Wikipedia_Shuffledtext100K<n<1M0 likes1.2k downloads4y agoHugging Face21rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes1k downloads10d agoHugging Face22agentlans /NousResearch-Hermes-3-Dataset-multiturn Hermes 3 Multiturn This is a filtered subset of NousResearch/Hermes-3-Dataset containing only multiturn conversations with more than three messages. Conversations with repetitive or trivial replies (for example, repeated "OK") have been excluded to improve quality. text10K<n<100K2 likes984 downloads1y agoHugging Face23BSC-LT /multi_lmentry Multi-LMentry This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?". Dataset Details Dataset Description Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.textquestion-answering100K<n<1M12 likes960 downloads5mo agoHugging Face24BelleGroup /multiturn_chat_0.8M Multiturn Chat 0.8M 内容 包含约80万条由BELLE项目生成的用户与助手的多轮对话。 注意:此数据集是由ChatGPT产生的,未经过严格校验,内容可能包含错误。使用过程中请注意这一点。 instruction中包含多轮对话的上文内容,以Human:和Assistant:区分,output中包含当前助手角色的回答。 样例 { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M.text100K<n<1M146 likes879 downloads3y agoHugging Face25facebook /multilokogated MultiLoKo: a multilingual local knowledge benchmark for LLMs MultiLoKo is a multilingual knowledge benchmark, covering 30 languages plus English. The questions are separately sourced for each language, with an annotation protocol designed to target locally relevant topics for the respective language. MultiLoKo contains the original data for each language, as well as both human and machine-authored translations of each non-English subset into English and vice versa, facilitating… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multiloko.textquestion-answering10K<n<100K7 likes861 downloads1y agoHugging Face26SetFit /amazon_reviews_multi_ja#amazon reviews multi japanese This dataset is a port of the official ['amazon_reviews_multi' dataset] (https://huggingface.co/datasets/amazon_reviews_multi) on the Hub. It has just the Japanese language version. It has been reduced to just 3 columns (and 4th "label_text") that are relevant to the SetFit task. text100K<n<1M7 likes800 downloads5y agoHugging Face27gonced8 /multi-session_chatNot my dataset, I only cleaned the dataset from ParlAI - Msc. text10K<n<100K8 likes796 downloads3y agoHugging Face28Multilingual-Multimodal-NLP /IfEvalCode-Instructtext1K<n<10K2 likes726 downloads1y agoHugging Face29Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes706 downloads1y agoHugging Face30eddie-OB /gsm8k-multilingual-reasoning gsm8k-multilingual-reasoning GSM8K with reasoning translated to multiple languages Schema {"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}} Usage from datasets importload_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K1 likes681 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.