datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.vehicle-mixed-traffic-detection
Visaitech Mixed-Traffic Vehicle Detection Dataset (v0.1)
Dashcam frames annotated for pedestrian / 2-wheeler / 3-wheeler / 4-wheeler
detection in South Asian mixed traffic, a class taxonomy general-purpose
COCO-trained detectors don't cover (COCO has no concept of an auto-rickshaw
or motorcycle-vs-bicycle-as-one-class "2-wheeler" grouping tuned for how
this traffic actually mixes on the road).
This is an early v0.1 release: 293 annotated frames from 6 source videos,
published… See the full description on the dataset page: https://huggingface.co/datasets/visaitech/vehicle-mixed-traffic-detection.MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.diffbir-mixed-setsLIVR_mixed
LIVR_mixed (v2)
Curated multimodal data for training a Qwen2.5-VL self-reflection RL pipeline,
plus the held-out LIVR splits and three external benchmarks (BLINK +
PixMo-Count + VSP) used to evaluate it.
Layout
train/ # LIVR train (9 tasks × 1000)
metadata.jsonl
livr_v2_manifest.json
images/<task>/... # ~8.7 GB
livr_eval/ # LIVR's own held-out val + test (8 tasks; counting → pixmo_count_eval)
validation/… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/LIVR_mixed.QariOCR-v0.3-markdown-mixed-dataset
QARI Markdown Mixed Dataset
📋 Dataset Summary
The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding.
This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition.
This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.Mixed_Closed_with_DescriptionDorayakiLin_parquet_move_cylinder_in_holes_mixed_v1_lerobotvqa-mixed
VQA Mixed (GQA + VizWiz + VQAv2 subset)
This dataset is a curated mixture of three major VQA benchmarks, subsampled and reformatted into Parquet shards for efficient model training (e.g., for VLMs like LLaVA or Qwen-VL).
Dataset Details
The dataset consists of image-question-answer triplets stored in Parquet format.
Data Splits
Train: train/data-*.parquet
Validation: validation/data-*.parquet
Columns
image: The visual input (Image feature).… See the full description on the dataset page: https://huggingface.co/datasets/kollessisopod/vqa-mixed.ur5_mixed_51_smoothThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5e",
"total_episodes": 51,
"total_frames": 41160,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LPSlvlv/ur5_mixed_51_smooth.MixBench2026
MixBench: A Benchmark for Mixed Modality Retrieval
MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs.
MixBench includes four subsets, each curated from a different data source:
MSCOCO
Google_WIT
VisualNews
OVEN
Each subset contains:
queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench2026.rt2rand_mixed_5tasks_clean50_randomized50
RoboTwin 2.0: clean50 + randomized50, five-task mixture
LeRobot v2.1 dataset for OpenPI training, converted from
TianxingChen/RoboTwin2.0.
Composition
The dataset contains 500 episodes and 111,025 frames:
Task
Embodiment
Clean
Randomized
turn_switch
aloha-agilex
50
50
move_playingcard_away
aloha-agilex
50
50
beat_block_hammer
aloha-agilex
50
50
place_dual_shoes
aloha-agilex
50
50
open_microwave
aloha-agilex
50
50
The clean episodes are all… See the full description on the dataset page: https://huggingface.co/datasets/s1ghhh/rt2rand_mixed_5tasks_clean50_randomized50.OpenGameArt-Mixed-Licenses
Dataset Card for OpenGameArt-Mixed-Licenses
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.geometry3k_4choices_mixedmhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.pdas-iter2-mixedmhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.mixedKhENTextLinereal_robot_data_arrangement_mixedmhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.rlbench_mixed_25_08_14_openpi_lerobotDMin_mixed_datasets_8846Implementation for "DMin: Scalable Training Data Influence Estimation for Diffusion Models".
Influence Function, Influence Estimation and Training Data Attribution for Diffusion Models.
Github, Paper
letzcross-wiki-mixed-en-fr-de
LëtzCross Wiki Mixed EN-FR-DE
Dataset Description
omarelba/letzcross-wiki-mixed-en-fr-de is a three-language visual-document-retrieval dataset. Every row contains one question in English, French, or German and a rendered image of the relevant PDF page. The question appears in query, while language records its language.
The dataset is used for multilingual late-interaction page-image retrieval training. The model receives query as the text input and image as the… See the full description on the dataset page: https://huggingface.co/datasets/omarelba/letzcross-wiki-mixed-en-fr-de.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.converted_mixed_pickandplace_datasetrealworld_replay_task820_firm_gentle_mixed_current
