datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MobileManiBenchr34-e621-7gb-public-2007-2014-oldhausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
lerobot_arnold_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 623,
"total_frames": 69062,
"total_tasks": 36,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:623"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/lerobot_arnold_data.arnoldThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 3570,
"total_frames": 394093,
"total_tasks": 458,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:3570"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paragon7060/arnold.lerobot_arnold_data_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2,
"total_frames": 201,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/lerobot_arnold_data_test.ARNOLD_multispace_3viewarnold-success-only
ARNOLD Success-Only
This is the success-only LeRobot v3.0 derivative of
paragon7060/arnold.
Episodes were retained only when the corresponding raw simulation rollout's
metadata.json had success: true. The filter removed 131 failed episodes
(25,530 frames) from the 3,570-episode source dataset.
Dataset summary
Field
Value
Episodes
3,439
Frames
368,563
Language tasks
454
FPS
10
Cameras
5
Resolution
480 x 480
Video codec
H.264
Episode… See the full description on the dataset page: https://huggingface.co/datasets/paragon7060/arnold-success-only.Bhagavad-Gita-Vyasa-Edwin-Arnold
Bhagavad Gita QA Dataset
Description
This dataset contains 500 question-answer pairs based on Edwin Arnold's translation of the Bhagavad Gita. The questions cover various aspects of the text, including philosophical concepts, characters, events, and teachings from the ancient Indian scripture.
Structure
The dataset is provided in two formats:
CSV format: bhagavad_gita_qa.csv
Parquet format: bhagavad_gita_qa.parquet
Each record contains two fields:
question:… See the full description on the dataset page: https://huggingface.co/datasets/sweatSmile/Bhagavad-Gita-Vyasa-Edwin-Arnold.arnold_data_pickup_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FrankaPanda",
"total_episodes": 623,
"total_frames": 69062,
"total_tasks": 36,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:623"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paragon7060/arnold_data_pickup_fixed.atlante-filiere-italiane
Atlante delle filiere italiane — edizione 2026
Dati aperti sulla struttura delle filiere produttive italiane: quante imprese, quanto
fatturano, quanti addetti impiegano e dove si concentrano.
Elaborazione Clientium su anagrafiche del Registro Imprese.
DOI: 10.5281/zenodo.21541218 · Pagina: https://clientium.it/atlante/
Licenza: CC BY 4.0 (riuso libero, anche commerciale, con citazione)
Aggiornamento dei dati: luglio 2026 · pubblicato il 2026-07-25
Come citare… See the full description on the dataset page: https://huggingface.co/datasets/Arnoldkoci/atlante-filiere-italiane.my-cool-datasetosservatorio-cold-email-italia
Osservatorio Cold Email Italia
Dati aperti sulla cold email B2B italiana, pubblicati da Clientium.
Licenza CC BY 4.0: riuso libero, anche commerciale, con citazione della fonte.
Repository ufficiale: github.com/arnold222a/osservatorio-cold-email-italia
Dataset Description
L'Osservatorio raccoglie due pubblicazioni di dati aperti:
1. Indice Outbound Italia (rilevazione trimestrale)
Benchmark sul reply rate della cold email B2B in Italia, calcolati… See the full description on the dataset page: https://huggingface.co/datasets/Arnoldkoci/osservatorio-cold-email-italia.execoconut-dataset
execoconut-dataset
ExecCoCoNuT (Execution CoCoNuT): Python code execution traces for continuous latent thought training.
Dataset Description
This dataset contains 1,000+ Python code snippets with execution traces, designed for training language models to reason about program state in continuous latent space.
Dataset Statistics
Samples: ~1,000 (train: 900, val: 50, test: 50)
Variable scope: 6 integer variables (a-f)
Operations: +, -, *, // (integer division)… See the full description on the dataset page: https://huggingface.co/datasets/ArnoldMoya/execoconut-dataset.dorothea_arnold_fireemblem
Dataset of dorothea_arnold (Fire Emblem)
This is the dataset of dorothea_arnold (Fire Emblem), containing 22 images and their tags.
The core tags of this character are brown_hair, green_eyes, long_hair, breasts, earrings, large_breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size
Download… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/dorothea_arnold_fireemblem.odia-tts-descriptionsindice-outbound-italia
Indice Outbound Italia — Q3 2026
Dati aperti sul reply rate della cold email B2B in Italia, pubblicati da Clientium.
DOI: 10.5281/zenodo.21541001 · Licenza: CC BY 4.0 · Periodicità: trimestrale
Per quanto ci risulta è il campione italiano più ampio pubblicato con metodologia dichiarata: i benchmark citati in Italia sono quasi sempre dati statunitensi.
I numeri
Calcolati sulle 556 campagne con almeno 500 invii, estratte da un dataset di 723 campagne cold email B2B… See the full description on the dataset page: https://huggingface.co/datasets/Arnoldkoci/indice-outbound-italia.ArnoldSchwarzenegger_rvc2pseudo-labeled-datasetlerobot_arnoldThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FrankaPanda",
"total_episodes": 3571,
"total_frames": 15426,
"total_tasks": 458,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:3571"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paragon7060/lerobot_arnold.my_datasetarnoldodia-2speaker-ttsKalkinArnoldArnoldAgnesCedric
