datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peoples_speech
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.mlcr-dataset
MLCR Dataset
Public dataset for the Multi-Length Context Retrieval evaluation experiments.
Configs
Config
Description
questions
Prompts, expected answers, and difficulty per case
cases_summaries
Document summaries for each case
filler_files
OCR text from filler pages
Usage
from datasets import load_dataset
ds = load_dataset("Wisedocs/mlcr-dataset", "questions")
print(ds["train"][0])
mlcs_en_pitchmlcMLCQA-dataset
MLCQA v1
Multilingual Language Chart Question Answering — A large-scale
multilingual Chart VQA dataset covering 11 languages and 8 chart types.
Dataset Summary
Metric
Count
Total Images
440,000
Total QA Pairs
4,400,000
Languages
11 (bn, en, gu, hi, kn, ml, mr, ne, pa, ta, te)
Chart Types
8 (area, donut, grouped_bar, horizontal_bar, line, pie, stacked_bar, vertical_bar)
Questions per Image
10
Random Seed
22
Splits
Split
Images… See the full description on the dataset page: https://huggingface.co/datasets/MLCQA/MLCQA-dataset.ML_Conferences-Peer-Reviews
Sem-Detect: ML Conference Peer-Review Authorship Dataset (ICML 2026)
This dataset contains over 22,000 peer reviews from ICLR and NeurIPS spanning three authorship classes: human-written, fully AI-generated, and LLM-refined (human reviews polished by an LLM).
It is the primary benchmark for training and evaluating Sem-Detect, an AI-Text Detection approach that combines textual features with claim-level semantic analysis, tailored for the peer-review domain.
Paper 📄… See the full description on the dataset page: https://huggingface.co/datasets/Sem-Detect/ML_Conferences-Peer-Reviews.mlcb
Dataset Card for "mlcb"
More Information needed
mteb-nl-vabb-mlcls-pr
VABBMultiLabelClassification
An MTEB dataset
Massive Text Embedding Benchmark
This dataset contains the fourteenth edition of the Flemish Academic Bibliography for the Social Sciences and Humanities (VABB-SHW), a database of academic publications from the social sciences and humanities authored by researchers affiliated to Flemish universities (more information). Publications in the database are used as one of the parameters of the Flemish performance-based research funding system… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-vabb-mlcls-pr.MLCID
Dataset Card for MLCID
Dataset Description
MLCID (Multi-layered Composable Image Dataset) is a high-quality dataset designed for text-guided multi-layered composable image synthesis.
The dataset includes detailed foreground and background layers, instance-level bounding boxes, and precise masks,
enabling advanced image synthesis and alignment learning between layers and text.
Uses
The mask can be read by the code below:
import pycocotools.mask as mask_util… See the full description on the dataset page: https://huggingface.co/datasets/huangrh9/MLCID.mlcs_fr_pitchMLC_instruction_25_langsfranka_actions_0215This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 64,
"total_frames": 7040,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:64"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0215.mlcs_et_pitchmlc-video-generation-datasetfranka_actions_0216_PandSThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 16,
"total_frames": 3456,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0216_PandS.franka_actions_0217_PandS_inte0.05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0217_PandS_inte0.05.record-stack-4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 31,
"total_frames": 12449,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/record-stack-4.record-pickplace-0310This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 12,
"total_frames": 3426,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 25,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/record-pickplace-0310.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 1255,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/record-test.0309-drop-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 32,
"total_frames": 5920,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/0309-drop-test.franka_actions_0217_PandS_inte0.02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0217_PandS_inte0.02.franka_actions_0220_ee_absThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 149,
"total_frames": 27714,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:149"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0220_ee_abs.mlcs_rw_pitchMLC_translated_11_langs_20240801record-pickplace-0310-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 76,
"total_frames": 21865,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 25,
"splits": {
"train": "0:76"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/record-pickplace-0310-merged.mlcourse_imdb_tripletmlcs_de_pitchfranka_actions_0223_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 66,
"total_frames": 12210,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:66"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0223_test.0223_act_policy_evalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 3,
"total_frames": 555,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/0223_act_policy_eval.franka_actions_0211This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 2048,
"total_frames": 286720,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:2048"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlcf-robot/franka_actions_0211.
