datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiviewX_Labelsafrica-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
SaleData
SalesLLM-10k (SaleData)
⭐ If you find this project helpful, please give us a star on GitHub! It means a lot to us.
The official data repository of SalesLLM: Benchmarking LLM Realistic Selling Skill — accepted by EMNLP 2026 as a Main Paper.
This repository hosts the SalesLLM-10k dataset: 10,000 high-quality, multi-turn sales conversations in Chinese across financial services (bank deposits, insurance, fund investment, stocks) and consumer products.
🔗 Related… See the full description on the dataset page: https://huggingface.co/datasets/MultiSense/SaleData.Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.MultiEdit
🧩 MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
📃 Arxiv
🚀 Dataset Overview
Based on our MLLM-driven data construction pipeline using GPT-4o and GPT-Image-1, we introduce MultiEdit, a comprehensive large-scale
instruction-based image editing dataset comprising over 107K samples targeting 6 challenging image editing tasks covering 56 subcategory
editing types (18 non-style-transfer and 38 style transfer). We also release… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/MultiEdit.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.Multi-Querier_Dialogue
Multi-Querier Dialogue (MQDialog) Dataset
📑 Dataset Details
Dataset Description
This is the dataset of "Querier-Aware LLM: Generating Personalized Responses to the Same Query from Different Queriers".
The Multi-Querier Dialogue (MQDialog) dataset is designed to facilitate research in querier-aware personalization. It contains dialogues with various queriers for each reponder. The dataset is derived from English and Chinese scripts of popular TV shows and… See the full description on the dataset page: https://huggingface.co/datasets/Nidhogg-zh/Multi-Querier_Dialogue.risale-nur-multilingual
Risale-i Nur Multilingual Corpus
Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir.
Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.sd-multiplayer-dataTo access an image use the following
Bucket URL: https://d26smi9133w0oo.cloudfront.net/
example:
https://d26smi9133w0oo.cloudfront.net/room-7/1670520485-CZk4C72xBr5wPfTpwDAnG6-7648_7008-a-chicken-breaking-through-a-mirrornnotn.webp
Bucket URL/key
SQLite
https://huggingface.co/datasets/huggingface-projects/sd-multiplayer-data/blob/main/rooms_data.db
sqlite> PRAGMA table_info(rooms_data);
0|id|INTEGER|1||1
1|room_id|TEXT|1||0
2|uuid|TEXT|1||0
3|x|INTEGER|1||0
4|y|INTEGER|1||0
5|prompt|TEXT|1||0… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-projects/sd-multiplayer-data.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.multi-legal-bench
Multi-Legal-Bench
Paper: Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries (v3)
Identical legal tasks evaluated on native court decisions from national registries in France, the Netherlands, Poland, the Czech Republic and Lithuania, with Ukrainian cells in the companion UA-Legal-Bench (not included here). Labels come from registry metadata. The v3 paper is an audit of what those labels let a benchmark measure: in… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/multi-legal-bench.lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private
Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b
Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.nfqa-multilingual-dataset
NFQA Multilingual Dataset
A large-scale multilingual dataset for Non-Factoid Question Answering (NFQA) classification, covering 49 languages and 8 question categories.
Dataset Statistics
Split
Examples
Train
28,653
Validation
3,539
Test
3,671
Total (Balanced)
35,863
Full Dataset (High Quality)
63,647
Dataset Composition
Languages (49 total)
Arabic (ar), Azerbaijani (az), Bulgarian (bg), Bengali (bn), Catalan (ca)… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/nfqa-multilingual-dataset.multiplexing-artiq-results
Multiplexing ARTIQ experimental results
This dataset is a public snapshot of the HDF5 result files from
artiq_working_dir - 20240315/results, collected for trapped-ion multiplexing
and quantum-networking experiments.
Contents
artiq-results.tar.zst: 7,433 HDF5 result files organized by date and hour.
MANIFEST.sha256: SHA-256 digest for every extracted HDF5 file.
DATASET_INFO.json: source inventory and publication audit summary.
The data span 170 date directories… See the full description on the dataset page: https://huggingface.co/datasets/bingran-you/multiplexing-artiq-results.multimodality-poc-llama31-ruler16k
Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K
Raw pre-RoPE query and hidden-state tensors captured during prefill, used
to study whether the per-(layer, kv_head) query distribution is unimodal
Gaussian (the assumption underpinning Expected Attention's MGF closed-form
in kvpress).
What's in here
65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5
prompts).
Each file (~414 MB) contains:
field
dtype
shape
meaning
hidden
float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.smash-karts-multiplayer-trajectory
Smash Karts 멀티플레이 구현해줘
A single Codex coding-agent session implementing a multiplayer browser-based 3D
kart battle game inspired by Smash Karts.
The request covers multiplayer play, game rules, weapons and effects, keyboard
controls, research, and implementation. The trajectory records the development
process, tool calls and results, validation work, and the final response.
Field
Value
Session title
Smash Karts 멀티플레이 구현해줘
Session ID… See the full description on the dataset page: https://huggingface.co/datasets/amsminn/smash-karts-multiplayer-trajectory.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.cs_squad-3.0
Dataset Card for Czech Simple Question Answering Dataset 3.0
This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section.
Dataset Description
The data contains questions and answers based on Czech wikipeadia articles.
Each question has an answer (or more) and a selected part of the context as the evidence.
A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private.Multilingual-BenchmarkThese are the GSM8K and ARC dataset translated by Google Translate.
BibTex
@misc{lu2024languagecountslearnunlearn,
title={Every Language Counts: Learn and Unlearn in Multilingual LLMs},
author={Taiming Lu and Philipp Koehn},
year={2024},
eprint={2406.13748},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2406.13748},
}
agent-spaces-tracesmultilingual-speech
Multilingual Indian Conversational Speech
A dataset of naturalistic, spontaneous two-speaker conversations across
13 Indian languages, with segment-level transcripts, speaker profiles,
timestamps, and recording metadata. Designed for ASR, TTS, speaker
diarization, and conversational speech research.
Languages (13)
Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali,
Odia, Punjabi, Tamil, Telugu, Urdu.
Content
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.exp186-sf-multipv-2m
exp186 Stockfish MultiPV soft targets (~2M)
Local Stockfish 18 full-strength MultiPV harvest for chess transformer soft-policy training.
Contents
data/positions_*.jsonl — ~2,000,000 positions (400 shards)
soft_cache.pt — training-ready tensor cache (compact move vocab 1968)
Labeling
Engine: Stockfish 18 (full strength, no UCI_LimitStrength)
MultiPV: 8
Depths: sampled in [2, 8] (weighted toward mid depths)
Soft probs: softmax(cp / τ) with τ=120… See the full description on the dataset page: https://huggingface.co/datasets/avewright/exp186-sf-multipv-2m.tasklist-grok-multilingual-100000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-org
Run Parameters
Parameter
Value
Model
grok-4-1-fast-reasoning
Temperature
0.9
Total Tasks
83052
Concurrency
30 workers
API Base
https://api.x.ai/v1
Generated
2026-04-07 14:31:14
Budget Cap
$15.0000
Multilingual
Yes (en, de, fr, es, nl, zh, ar, ru)
Language Distribution
Language
Code
Tasks
Arabic
ar
10446
German
de
10397
Dutch
nl
10353
Spanish
es
10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.lm-eval-results-MTSAIR-multi_verse_model-private
Dataset Card for Evaluation run of MTSAIR/multi_verse_model
Dataset automatically created during the evaluation run of model MTSAIR/multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MTSAIR-multi_verse_model-private.openorca-multiplechoice-10kA 10k subset of OpenOrca dataset, focusing on multiple choice questions.
Credit to Tian Xia.
