datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-SWE-RL-Verified
Multi-SWE-RL-Verified
Gold-patch-validated subset of
PrimeIntellect/Multi-SWE-RL-Reupload
(ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across
C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end.
Default dataset of the multiswe_v1 taskset.
Changes vs upstream
Starting from the 4,703-row re-upload:
C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are
broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
MT-Reasoning
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses.
lang
rows
prompt_tokens
reasoning_tokens
response_tokens
total_tokens
deu_Latn
17_354_716
1_873_153_732
26_010_932_738
14_862_651_336
42_746_737_806
fra_Latn
17_354_716
1_802_885_115
25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.knowchat-multi-turn-dialogues
KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks
KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.risale-nur-multilingual
Risale-i Nur Multilingual Corpus
Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir.
Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.SWE-ZERO-multilang-300-trajectories
SWE-ZERO-multilang-300-trajectories
300 execution-free agentic rollouts from
ricdomolm/mini-coder-1.7b
across all 20 programming languages in
nebius/SWE-rebench-V2
(5 PRs per language × 3 rollouts per PR), generated as part of
marin-community/marin#4653.
Companion to the Python-only run:
AlienKevin/SWE-ZERO-1k-trajectories-32k — 1,000 rollouts on 100 Python PRs (10 repos × 10 PRs × 10 rollouts)
AlienKevin/SWE-ZERO-1k-trajectories— the original 8k-context Python baseline
Each row… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/SWE-ZERO-multilang-300-trajectories.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.novel-multilingual
WebNovel Multilingual Dataset
Dataset Description
This dataset contains web novels scraped from WebNovel.com across multiple languages. Each entry includes the complete novel content with chapter information, metadata, and classification tags.
Note: This dataset excludes content in the following languages: id, ID
Dataset Statistics
Total Novels: 8,324
Total Chapters: 233,410
Total Characters: 1,617,589,129
Languages: 10
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-multilingual.multilingual-textarena-ColonelBlotto-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-ColonelBlotto-v0-train.ultrafeedback-multi-binarized-preferences-cleaned
UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned,
and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the
AlpacaEval, MT-Bench, and LM Eval Harness benchmarks.
Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.multilingual-textarena-SimpleTak-v0-train-v2
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.multilingual-reasoning-gym-sft
Reasoning Gym SFT Dataset
This dataset contains Supervised Fine-Tuning (SFT) reasoning data procedurally generated using Reasoning Gym environments.
It is designed to train reasoning models (such as DeepSeek-R1-style or Qwen-Coder-style models) to explain their step-by-step reasoning chain before outputting a final answer wrapped inside LaTeX \boxed{...}.
Where Does This Dataset Come From?
This dataset is procedurally generated from Reasoning Gym, an open-source… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/multilingual-reasoning-gym-sft.multimodel-capitulation-interp
Multi-model wrongful-capitulation internal-readout dataset
Per-turn internal readouts + behavioral labels from two-model collaborative conversations (Qwen2.5-3B-Instruct × gemma-2-2b-it) on 6 reasoning benchmarks, restricted to the disagreement subset (one model right, one wrong solo). Built to test whether a linear correctness probe on the residual stream can predict wrongful capitulation (a model abandoning an answer it knew was correct under a partner's wrong assertion)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/multimodel-capitulation-interp.Uganda-Multilingual-QA
BUAIR Uganda Multilingual Q&A
Parallel question–answer dataset for Ugandan languages, curated by the BUAIR Voice initiative at Busitema University.
Dataset version: v2 (updated 2026-08-31)
Each language config contains the same 4,256 agriculture / rural-livelihood Q&A pairs, with English as the shared source and translations into Japadhola, Ateso, Runyankore, and Luganda.
Changelog (v2)
Replaced v1 data (5,000 pairs from Multiligual-QA.xlsx) with cleaned data… See the full description on the dataset page: https://huggingface.co/datasets/BUAIR/Uganda-Multilingual-QA.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.multilingual-textarena-ColonelBlotto-v0-train-v2
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-ColonelBlotto-v0-train-v2.multilingual-textarena-KuhnPoker-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-KuhnPoker-v0-train.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.multilingual-textarena-SimpleTak-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.csdi-multilingual
CSDI: A Fine-Grained Fundus Image Dataset of Cataract Severity and Diagnostic Images
Acknowledgment: This dataset is a multilingual extension of the original CSDI Cataract Diagnosis Dataset created by Xie, Z., Ao, M., Tang, H. et al. The original dataset provides expert written reports and diagnostic descriptions in English and Chinese. This version expands the diagnostic text into 31 additional languages to support cross-lingual research in automated cataract screening and… See the full description on the dataset page: https://huggingface.co/datasets/projetogabi/csdi-multilingual.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.multilingual-textarena-Nim-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset = load_dataset("The-CoLab/multilingual-textarena-Nim-v0-train"… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-Nim-v0-train.MultimodalUnlearningEvalBenchmark
🧠 Multimodal Unlearning Evaluation Benchmark
📌 Overview
This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning.
It supports reproducibility of results in:
Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026)
📊 Contents
File
Description
📄 multimodal_results.json
Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench)
📄 unimodal_results.json
CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.multilingual-textarena-TicTacToe-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-TicTacToe-v0-train.Multi_CodeNet4Repairmultilingual-textarena-Nim-v0-train-v2
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-Nim-v0-train-v2.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.
