datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pose6daug
pose6daug
Real-world Franka manipulation episodes with object-swap and action augmentation
artifacts. 120 training episodes over 4 objects (blue_cup, green_pear, kanu,
white_spray), dual ZED cameras (exo static + ego wrist-mounted).
Layout
Per-frame PNGs are packed into uncompressed tars per episode — the dataset has
~427k mask/plate frames and loose files hit Hugging Face's per-repo file
recommendation and API rate limits hard.
data/<object>/<NNNN>/
masks.tar… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/pose6daug.reward-projection-goal-generalisation-vlmgoat
Dataset Card for Dataset Name
Dataset Summary
The dataset.json file contains ~1.7 million synthetic data for arithmetic tasks, generated by dataset.ipynb.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/goat.goai_2026_lerobot_realpi05-libero-goal-task-8-evidence
Robium Pi0.5 LIBERO-Goal Task 8 evidence
This is the public evidence bundle for Robium issue #69. It records one fixed,
no-retry evaluation of lerobot/pi05_libero_finetuned_v044 on LIBERO-Goal task
8, put_the_bowl_on_the_plate, using the canonical prompt “put the bowl on the
plate.”
Result
20/20 successful episodes; the predeclared target was 16/20.
Fixed initial states 0–19 map to seeds 1000–1019.
Batch size 1, hard environment/policy reset before every episode… See the full description on the dataset page: https://huggingface.co/datasets/robium/pi05-libero-goal-task-8-evidence.libero_plus_goal
libero_plus_goal: detailed LeRobot v3.0
This dataset was converted from the LIBERO Plus LeRobot v2.1 libero_plus_goal partition.
The original 8D state and 7D action vectors are preserved exactly as
raw_state.ref_state and raw_action.ref_action. Canonical low-dimensional fields follow
failure_rollout_data/dataset.md; debug.gripper_eef_* contains the ground-truth next-step
relative EEF motion for inspection.
Required camera transform for canonical training
The… See the full description on the dataset page: https://huggingface.co/datasets/typoverflow/libero_plus_goal.goatdataSustainable_Development_Goals_QA_V2
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
Sustainable_Development_Goals_QA
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
libero_goal_mj332
libero_goal_no_noops_lerobot: detailed LeRobot v3.0
This dataset was converted from the Fast-WAM LIBERO MuJoCo 3.3.2 LeRobot v2.1
libero_goal_no_noops_lerobot partition. It contains successful demonstrations whose historical no-op actions
were removed by simulator replay before this conversion. This converter preserves all remaining
frames and does not apply any additional filtering.
The original 8D state and 7D action vectors are preserved exactly as
raw_state.ref_state and… See the full description on the dataset page: https://huggingface.co/datasets/typoverflow/libero_goal_mj332.goal-pashto-chat-sharegpt-5GB
📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset
A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment.This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge.
📌 Dataset Summary
goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB.nepi-prompts-dataset
NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized)
Dataset Summary
This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs).It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts.
Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav-GOAT/nepi-prompts-dataset.databird-goalsGoan_Datapsyllmgoat-mptgo_ablationGoalman12goatis-transcripts
Goatis / Sv3rige Video Transcripts
Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the
YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the
current Goatis channel (2019–2026). This is the dataset behind
goatis.net, a searchable archive in the style of
aajonus.net.
What makes it more than raw ASR
Every video was processed with speaker identification, not just
transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.kobza-2m-jsonlgoat-chinesegoat中文算术数据集
将goat数据集的Template,更换成中文的Template,数学表达式不变
Swift-libero-goal-Actiongoai-bench-resultshubert-baseGoalman06OpenOrca-solution-for-a-goal-viGoalman03Goalman14Goan_churchesLibero-Goal-subtask
