datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiArithVript_Multilingual
🎬 Vript: A Video Is Worth Thousands of Words [Github Repo]
We construct another fine-grained video-text dataset with 19.1K annotated high-resolution UGC videos (~677k clips) in multiple languages to be the Vript_Multilingual.
New in Vript_Multilingual:
Multilingual: zh (60%), en (17%), de (15%), ja (6%), ko (2%), ru (<1%), es (<1%), pt (<1%), jv (<1%), fr (<1%), id (<1%), vi (<1%)
More diverse and fine-grained categories: 113 categories (please check vript_CN-V2_meta.json)… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Multilingual.IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.MultiHopRAG
Dataset Card for Dataset Name
A Dataset for Evaluating Retrieval-Augmented Generation Across Documents
Dataset Description
MultiHop-RAG: a QA dataset to evaluate retrieval and reasoning across documents with metadata in the RAG pipelines. It contains 2556 queries, with evidence for each query distributed across 2 to 4 documents. The queries also involve document metadata, reflecting complex scenarios commonly found in real-world RAG applications.
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/yixuantt/MultiHopRAG.clevr-multichange
CLEVR-Multi-Change (30–40 objects)
Two-image change-captioning data used in "Stateful Visual Encoders for
Vision-Language Models" (the Multi-object Visual Differencing task). Each example is a before/after pair of a CLEVR scene
with 30–40 objects and 4 simultaneous changes (add / delete / move /
replace), rendered at 768×768 with a wide camera angle. Built with the
CLEVR-Multi-Change engine (Johnson et al. 2017; Qiu et al. 2021).
Code & paper:… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/clevr-multichange.McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
Multi-SWE-bench
SWE-bench-Java: A GitHub Issue Resolving Benchmark for Java
📰 News
[Aug. 27, 2024]:We’ve released the JAVA version of SWE-bench! Check it out on Hugging Face. For more details, see our paper!
📄 Abstract
GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/Daoguang/Multi-SWE-bench.IfEvalCode-testsetamazon_reviews_multi_enmultilingual_mbppMBPP translated to 15 programming languages using o4-mini-medium.
source_language = "python"
target_languages = [
"cpp",
"c",
"javascript",
"java",
"php",
"csharp",
"typescript",
"bash",
"swift",
"go",
"rust",
"ruby",
"r",
"matlab",
"scala",
"haskell"
]
effort = "medium"
dataset_name = "google-research-datasets/mbpp"
model = "o4-mini"
africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.Nemotron-SFT-Multilingual-v2
Dataset Description:
Nemotron-SFT-Multilingual-v2 is a multilingual supervised fine-tuning (SFT) dataset for post-training text-generation models. It is generated by translating seed data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1, adding multilingual coverage for Hindi (hi), Korean (ko), Brazilian Portuguese (pt-br), and refreshed Japanese (ja) data.
The dataset is generated with a new data processing pipeline that avoids line-breaking… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v2.multi30k
Multi30k
This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here.
Each example consists of an "en" and a "de" feature. "en" is an English sentence, and "de" is the German translation of the English sentence.
Data Splits
The Multi30k dataset has 3 splits: train, validation, and test.
Dataset Split
Number of Instances in Split
Train
29,000
Validation
1,014
Test
1,000
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/bentrevett/multi30k.Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course).
This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.CADBench-Extended-Multimodal-Dataset
Dataset Card
Dataset Description
CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics.
Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.multinerd
Dataset Card for MultiNERD dataset
Description
Summary: In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/multinerd.MultipanelVQAMultiLegalPile_Wikipedia_Shuffledmulti-agent-coordination-transcripts
Multi Agent Coordination Transcripts
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.NousResearch-Hermes-3-Dataset-multiturn
Hermes 3 Multiturn
This is a filtered subset of NousResearch/Hermes-3-Dataset
containing only multiturn conversations with more than three messages.
Conversations with repetitive or trivial replies (for example, repeated "OK") have been excluded to improve quality.
multi_lmentry
Multi-LMentry
This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?".
Dataset Details
Dataset Description
Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.multiturn_chat_0.8M
Multiturn Chat 0.8M
内容
包含约80万条由BELLE项目生成的用户与助手的多轮对话。
注意:此数据集是由ChatGPT产生的,未经过严格校验,内容可能包含错误。使用过程中请注意这一点。
instruction中包含多轮对话的上文内容,以Human:和Assistant:区分,output中包含当前助手角色的回答。
样例
{
"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M.multiloko
MultiLoKo: a multilingual local knowledge benchmark for LLMs
MultiLoKo is a multilingual knowledge benchmark, covering 30 languages plus English.
The questions are separately sourced for each language, with an annotation protocol designed to target locally relevant topics for the respective language.
MultiLoKo contains the original data for each language, as well as both human and machine-authored translations of each non-English subset into English and vice versa, facilitating… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multiloko.amazon_reviews_multi_ja#amazon reviews multi japanese
This dataset is a port of the official ['amazon_reviews_multi' dataset] (https://huggingface.co/datasets/amazon_reviews_multi) on the Hub. It has just the Japanese language version. It has been reduced to just 3 columns (and 4th "label_text") that are relevant to the SetFit task.
multi-session_chatNot my dataset, I only cleaned the dataset from ParlAI - Msc.
IfEvalCode-InstructDEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
gsm8k-multilingual-reasoning
gsm8k-multilingual-reasoning
GSM8K with reasoning translated to multiple languages
Schema
{"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}}
Usage
from datasets importload_dataset
ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning")
print(ds["train"][0])
Source
Derived from OpenAI GSM8K.
MulDimIF
[ACL 2026] MulDimIF
A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
Data and code for the paper A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
Junjie Ye
jjye23@m.fudan.edu.cn
May. 13, 2025
Introduction
Instruction following refers to the ability of large language models (LLMs) to generate outputs that satisfy… See the full description on the dataset page: https://huggingface.co/datasets/Junjie-Ye/MulDimIF.
