datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
S10-Citadel-Core
Run and deploy your AI Studio app
This contains everything you need to run your app locally.
Run Locally
Prerequisites: Node.js
Install dependencies:
npm install
Set the GEMINI_API_KEY in .env.local to your Gemini API key
Run the app:
npm run dev
pose6daug
pose6daug
Real-world Franka manipulation episodes with object-swap and action augmentation
artifacts. 120 training episodes over 4 objects (blue_cup, green_pear, kanu,
white_spray), dual ZED cameras (exo static + ego wrist-mounted).
Layout
Per-frame PNGs are packed into uncompressed tars per episode — the dataset has
~427k mask/plate frames and loose files hit Hugging Face's per-repo file
recommendation and API rate limits hard.
data/<object>/<NNNN>/
masks.tar… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/pose6daug.mv-mesh-40kscripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.goat
Dataset Card for Dataset Name
Dataset Summary
The dataset.json file contains ~1.7 million synthetic data for arithmetic tasks, generated by dataset.ipynb.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/goat.kobzaKobza
On the Path to Make Ukrainian a High-Resource Language [paper]
Kobza is the largest publicly available Ukrainian corpus to date, comprising nearly 60 billion tokens across 97 million documents. It is designed to support pretraining and fine-tuning of large language models (LLMs) in Ukrainian, as well as multilingual settings where Ukrainian is underrepresented.
🧾 Dataset Summary
Kobza aggregates high-quality Ukrainian text from a wide range of web sources and applies… See the full description on the dataset page: https://huggingface.co/datasets/Goader/kobza.goan-konkani-speech
Goan Konkani Speech (Romi)
47,365 audio clips, 108.5 hours of Goan Konkani (ISO 639-3 gom)
speech from Goan television news, transcribed in Romi Konkani - Konkani written in
the Roman script.
Konkani is a low-resource language with very little public speech data. This is
assembled from broadcast news, so it is real spoken Konkani: studio anchors, field
reporters, phone interviews, and the Konkani-English code-switching that Goan speakers
actually use.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-speech.greek_sheep_goats_dataset
AGRARIAN-EU: Greek Sheep and Goats (GSG) Object Detection Dataset
This dataset was developed within the framework of the AGRARIAN-EU project.
It aims to provide high-quality training data for the detection of sheep and goats in realistic, diverse, and challenging agricultural environments.
📊 Dataset Structure
The dataset comprises 3,600 images (640x640 resolution), generated by partitioning 450 high-definition (1920x1080) video frames into 8 partially overlapping… See the full description on the dataset page: https://huggingface.co/datasets/AGRARIAN/greek_sheep_goats_dataset.goal-step-wikihowhttps://github.com/zharry29/wikihow-goal-step
@inproceedings{zhang-etal-2020-reasoning,
title = "Reasoning about Goals, Steps, and Temporal Ordering with {W}iki{H}ow",
author = "Zhang, Li and
Lyu, Qing and
Callison-Burch, Chris",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/goal-step-wikihow.lego_flat_final_goal_200_v1
LEGO flat final-goal dataset
200 robot demonstrations: 100 human-teleoperated episodes + 100 correction episodes, 79,956 frames at 20 Hz. This is simulated robot data; it is not the synthetic human-hand proxy dataset.
Use the visual_inspection viewer to browse every episode: front camera at start, middle and end, middle wrist camera, final goal image, and its text description. Full front and wrist MP4 recordings are in videos/.
Training representation
The fixed… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_flat_final_goal_200_v1.Goal-Dataset_en_infootball-charts-results-goal-timing
Football Charts — Match Results and Goal Timing
Football Charts publishes results, fixtures, league tables and the minute of
every goal for 93 leagues in 42 countries, including the lower divisions and
women's competitions most sources skip. Free JSON API and MCP server; the
results and goal-timing dataset is CC BY 4.0 with a DOI
(10.5281/zenodo.22295583).
Most public football datasets cover the big five European leagues and stop at
the final score. This one reaches Serie C, 3.… See the full description on the dataset page: https://huggingface.co/datasets/Damir81/football-charts-results-goal-timing.goai_2026_lerobot_realgoat
🐐 GOAT
Generalized Occupational Aptitude Test (GOAT) is a dataset based on questions from Russian government exams that is required for every person graduated from school.
Currently, the dataset cover questions from Literature, Sociology and Russian language subjects.
All questions are divided by expected output format:
Single choice. In such tasks, there is a set of possible answers from which you need to choose the right one. The answer to such tasks is one digit, which is the… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/goat.libero_goal_force_openvla_predicted_action_fixedgold-goat1B-1.5M-gens-4-30lt-vlm
lt-vlm
Image dataset with captions and metadata. Generated and maintained via Databl.ai.
Structure
images/ — image files
metadata.csv — per-image metadata. The file_name column links each row to its image under images/. Additional columns (e.g. caption, caption_verified, caption_translated) hold AI-generated and human-verified annotations.
Maintained with Databl.ai.
PORukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.turkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
dynavla-libero-goal-t7-frictionloss-vanilla-valThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 20,
"total_frames": 2305,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-vanilla-val.dynavla-libero-goal-t7-frictionloss-oracle-stall-unseenThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 20,
"total_frames": 2369,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-oracle-stall-unseen.libero-goal-points-patchpolicy
LIBERO-Goal point clouds for patch_policy_cleanup
The generated-not-downloaded files needed to train the point-cloud and point-caption conditioners in
llm-robotics/patch_policy_cleanup, plus the two
HDF5 trees they were derived from.
The base LIBERO demonstrations (images, actions, states) are not here — they come from
gaoyuezhou/patch-policy-datasets.
This dataset only supplies the per-demo files that repo's scripts generate, so a new machine does not have
to re-run them.… See the full description on the dataset page: https://huggingface.co/datasets/ParsaSharifi/libero-goal-points-patchpolicy.dynavla-libero-goal-t7-frictionloss-vanilla-trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 160,
"total_frames": 21354,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:160"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-vanilla-train.dynavla-libero-goal-t7-frictionloss-oracle-stall-trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 160,
"total_frames": 18035,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:160"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-oracle-stall-train.dynavla-libero-goal-t7-frictionloss-oracle-stall-valThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 20,
"total_frames": 2223,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-oracle-stall-val.pi05-libero-goal-task-8-evidence
Robium Pi0.5 LIBERO-Goal Task 8 evidence
This is the public evidence bundle for Robium issue #69. It records one fixed,
no-retry evaluation of lerobot/pi05_libero_finetuned_v044 on LIBERO-Goal task
8, put_the_bowl_on_the_plate, using the canonical prompt “put the bowl on the
plate.”
Result
20/20 successful episodes; the predeclared target was 16/20.
Fixed initial states 0–19 map to seeds 1000–1019.
Batch size 1, hard environment/policy reset before every episode… See the full description on the dataset page: https://huggingface.co/datasets/robium/pi05-libero-goal-task-8-evidence.dynavla-libero-goal-t7-frictionloss-oracle-gated-trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 160,
"total_frames": 12720,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:160"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-oracle-gated-train.dynavla-libero-goal-t7-frictionloss-oracle-gated-valThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 20,
"total_frames": 1544,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-oracle-gated-val.dynavla-libero-goal-t7-frictionloss-vanilla-unseenThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 20,
"total_frames": 2683,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vpraise00/dynavla-libero-goal-t7-frictionloss-vanilla-unseen.
