datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bagpiper_PreTrain_Data
Bagpiper Pretraining Data
Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated
with Bagpiper, an open-ended audio language
model that learns bidirectional mappings between audio and comprehensive text
descriptions across speech, music, environmental sound, and mixtures.
The en metadata describes the primary rich-caption language. Source audio can
contain speech or singing in other languages; it is not an English-only audio
guarantee.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.LichessGamesgo2-air-controlbench-v1
Go2 Air ControlBench v1 Public Preview
Go2 Air ControlBench is a compact command-to-outcome benchmark for a stock Unitree Go2 Air. It asks a narrow question that matters for robot planners and world-model scorers:
If we command the robot to move, what actually happens, and which candidate command should a planner have selected?
The release is deliberately small and claim-bounded. It is not an imitation-learning corpus and not mocap-grade ground truth. It is a public-safe… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/go2-air-controlbench-v1.so101-can-butler
SO-101 Can Butler
Teleoperated demonstrations of a human-triggered can handover on a low-cost SO-101 arm: the robot stays still until a person's hand appears on the mat, then reaches, grasps a can, and hands it over.
A SmolVLA fine-tune on this data — 20k steps, about $2.30 of rented RTX 4090 — grasped and delivered the can autonomously, verified 3 times out of 3 attempts. Model: espejelomar/smolvla-so101-can-butler.
What it looks like
A policy trained on… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/so101-can-butler.espeech_balalaika
ESpeech datasets (w/o podcasts) Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/espeech_balalaika.esp32-arm-test2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"base.pos",
"shoulder.pos",
"elbow.pos",
"wrist.pos",
"gripper.pos"
],
"shape": [
5
]
},
"observation.state": {… See the full description on the dataset page: https://huggingface.co/datasets/hieu24/esp32-arm-test2.aloha_sim_insertion_espada_finalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/J-joon/aloha_sim_insertion_espada_final.yam-espressoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "yam_follower",
"total_episodes": 91,
"total_frames": 72048,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:91"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robots123/yam-espresso.vulnerabilidades-ia-espanol
Vulnerabilidades CVE en sistemas de IA (espanol)
Corpus de advisories CVE/GHSA que afectan a paquetes y SDKs de IA (langchain, openai, anthropic, llamaindex, etc.) traducido al espanol por LaAutopsIA (ApisDom Intelligence Group). Fuente principal: GitHub Advisory Database (CC-BY-4.0).
Cifras del snapshot actual
14 vulnerabilidades publicadas en este snapshot.
Mes archivado: 2026-08.
Ultima edicion: 2026-09-01T21:35:37.468Z.
Frecuencia: sincronizacion mensual.… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/vulnerabilidades-ia-espanol.pinga-fogo-chico-xavier
🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971
As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela
TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp.
345 turnos (115 deles respostas do próprio Chico Xavier), a partir de
6 horas de áudio — o registro mais extenso do médium falando de improviso,
sem edição, diante de um painel de jornalistas.
Arquivos
Arquivo
Programa
Turnos
Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.indice-fallos-ia-espanol
Indice de Fallos IA en espanol
Snapshots mensuales del Indice de Fallos IA producido por el observatorio La AutopsIA (ApisDom Intelligence Group). Mide la fiabilidad de modelos LLM con benchmarks oficiales independientes, en formato citable y trazable.
Cifras del snapshot actual
765 mediciones en este snapshot.
Mes archivado: 2026-09.
Recomputado: 2026-09-01T21:34:26.058Z.
Frecuencia: sincronizacion mensual.
Para que sirve este dataset
Datos… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/indice-fallos-ia-espanol.ESP_DATA_TWOaloha_sim_transfer_cube_espadaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 6600,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/J-joon/aloha_sim_transfer_cube_espada.aloha_sim_insertion_espadaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 11350,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/file-{episode_chunk:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/J-joon/aloha_sim_insertion_espada.esp32s3-csi-har-2025
ESP32-S3 WiFi CSI human activity dataset (July 2025)
Labelled WiFi Channel State Information captured from an ESP32-S3 on
29 July 2025. Four postures. 80 recordings. Collected and released by
Bl4ckd09.
The capture
Property
Value
Radio
ESP32-S3, single board, Espressif CSI parser
Date
29 July 2025, 18:23 to 19:55 local
Sample rate
19.823 Hz mean across all 80 files (range 19.12 to 20.00)
Subcarriers
192, stored as 384 interleaved I/Q values… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/esp32s3-csi-har-2025.cell1_20260513_electronics-packing_esp-cable-2bb-jumpersmf-ff-dcy-dc_unpackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "starpilot_yam_gripper",
"total_episodes": 8,
"total_frames": 14466,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/starpilot-ai/cell1_20260513_electronics-packing_esp-cable-2bb-jumpersmf-ff-dcy-dc_unpacking.jatzingueni-purepecha-espaniolespresso-v2-carbon-water-data
ESPResso V2: Textile Carbon & Water Footprint Training Data
50,000 synthetic records for product-level carbon and water footprint prediction in textiles, generated by a 7-layer LLM-orchestrated pipeline with deterministic C99 calculation engines. Developed at the University of Amsterdam.
Dataset Description
The textile industry faces mounting regulatory pressure under the EU ESPR and Digital Product Passport mandate to quantify product-level environmental footprints.… See the full description on the dataset page: https://huggingface.co/datasets/Tr4m0ryp/espresso-v2-carbon-water-data.kising_score_segmentscell1_20260512_electronics-packing_esp-cable-2servoblue-bb-jumpersff-mm-prototypeboarThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "starpilot_yam_gripper",
"total_episodes": 5,
"total_frames": 15396,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/starpilot-ai/cell1_20260512_electronics-packing_esp-cable-2servoblue-bb-jumpersff-mm-prototypeboar.esperanto-boolq-questions
esperanto-boolq-questions
BoolQ questions (train + validation, 12,697 rows) translated from
English to Esperanto by
jensjepsen/eo-mt-v13-large-bidir,
with round-trip quality metadata for filtering.
Row schema
field
description
orig_idx
original BoolQ row index (train first, then validation)
split
source split (train / validation)
en_orig
raw BoolQ question (lowercase, no ?, as in google/boolq)
en_preproc
preprocessed input fed to MT: spaCy… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-boolq-questions.spain-nhs-waiting-lists-2025-12
ES·pera — Spain NHS Waiting Lists, December 2025
This dataset contains the December 2025 interannual extract published by ES·pera from Spain's official SISLE-SNS waiting-list data. It covers national and autonomous-community figures, selected first-consultation specialties, surgical specialties and procedures, with corresponding December 2024 comparison fields where published.
ES·pera is an independent Spanish public-data project that integrates and normalises official public… See the full description on the dataset page: https://huggingface.co/datasets/esperaorg/spain-nhs-waiting-lists-2025-12.espada_aloha_insertionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 50337,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 50,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/J-joon/espada_aloha_insertion.Fashion-MNIST-CSVThis dataset is a direct copy of Fashion-MNIST, originally published by Zalando Research on Kaggle https://www.kaggle.com/datasets/zalando-research/fashionmnist.
Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. Zalando intends Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for… See the full description on the dataset page: https://huggingface.co/datasets/vincent-espitalier/Fashion-MNIST-CSV.esperanto-metamath-gsmsanciones-laborales-espana
Sanciones por incumplimiento laboral en pymes (España)
Tabla estructurada de las sanciones asociadas a los incumplimientos laborales más habituales en pymes españolas: registro horario y jornada (LISOS), canal de denuncias (Ley 2/2023) y contrato a tiempo parcial (Estatuto de los Trabajadores). Cada fila indica el sujeto responsable, la gravedad, la horquilla de importes, la norma y el artículo concreto.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/sanciones-laborales-espana.cleaning-purp-cube-20-hard-espThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 25,
"total_frames": 11250,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jomell310/cleaning-purp-cube-20-hard-esp.cleaning-purp-cube-45-espThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 45,
"total_frames": 20250,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jomell310/cleaning-purp-cube-45-esp.espada_aloha_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 36569,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 50,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/J-joon/espada_aloha_cube.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8575,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/espressobot/record-test.
