datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fragbench
FragBench (Public Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Dataset Summary
FragBench is a benchmark for evaluating cross-session, fragmented attacks on
LLM agents that use tools via the Model Context Protocol (MCP). Each campaign
is decomposed into many small fragments distributed across sessions; a
defender must reconstruct the compositional intent. The public tier in this
repository… See the full description on the dataset page: https://huggingface.co/datasets/LidaSafety/fragbench.dual-lidar-umi-relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-relative.R1_Lite_cover_the_pot_lid
R1_Lite_cover_the_pot_lid
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value
Total Episodes… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_cover_the_pot_lid.dual-lidar-umiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi.stack_md_lidA copy of the deduplicated the stack Markdown files, annotated with fastText to include language labels (lid column) and probabilities (lid_prob). This was done using the NLLB fastText model facebook/fasttext-language-identification.
The following languages were detected:
abk_Cyrl
ace_Arab
ace_Latn
ady_Cyrl
afr_Latn
aka_Latn
als_Latn
amh_Ethi
arb_Arab
arb_Latn
arn_Latn
asm_Beng
ast_Latn
ayr_Latn
azb_Arab
azj_Latn
bak_Cyrl
bam_Latn
ban_Latn
bel_Cyrl
bem_Latn
ben_Beng
bho_Deva
bis_Latn
bjn_Arab… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/stack_md_lid.dual-lidar-combined-filtered-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-long-gripper.robot-umi-lidar3d-e2eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
1200,
1920,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/robot-umi-lidar3d-e2e.dual-lidar-combined-filtered
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 157 demonstrations (156492 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered.record-drop-lidThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 38,
"total_frames": 20653,
"total_tasks": 1,
"total_videos": 76,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:38"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rohanc007/record-drop-lid.dual-lidar-combined-filtered-joint-positions
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 157 demonstrations (156492 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains left YAM joints 0–5, normalized left gripper, right YAM joints 0–5, and normalized right gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved;… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-joint-positions.yapay_zeka_turkce_mmlu_liderlik_tablosu
Yapay Zeka Türkçe MMLU Liderlik Tablosu
Bu veri seti serisi, Türkiye’deki eğitim sisteminde kullanılan gerçek sorularla yapay zeka modellerinin Türkçedeki yeteneklerini değerlendirmeyi amaçlar. Çeşitli büyük dil modellerinin (LLM) Türkçe Massive Multitask Language Understanding (MMLU) benchmark'ı üzerindeki performansını değerlendirir ve sıralar. Bu veri seti, modellerin Türkçe anlama ve cevaplama yeteneklerini karşılaştırmak için kapsamlı bir bakış açısı sunar. Her modelin… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/yapay_zeka_turkce_mmlu_liderlik_tablosu.agilex_take_off_lid_of_ice_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 9285,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_take_off_lid_of_ice_box.dual-lidar-combined-filtered-joint-positions-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains left YAM joints 0–5, normalized left gripper, right YAM joints 0–5, and normalized right gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved;… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-joint-positions-long-gripper.close_laptop_lid_25_08_03_lerobotv21egostation-iphone-lidar-household-v1
Zen-O Household Manipulation, iPhone LiDAR
8 first-person recordings of ordinary household work, with both hands
tracked in three dimensions and in real metres, laid out in LeRobot v2.1.
Episodes
8
Frames
128,833 at 30 fps, about 71.6 minutes
Video
observation.images.head, 1920x1440
State
7 floats, camera position and orientation
Action
20 floats, both wrists and both grippers
Coordinate frame
ROS REP 103, X forward, Y left, Z up, metric… See the full description on the dataset page: https://huggingface.co/datasets/zeno-labs/egostation-iphone-lidar-household-v1.close_laptop_lid_1000_25_08_09_lerobotv21pick_up_lidThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 30,
"total_frames": 25564,
"total_tasks": 2,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pick_up_lid.robot-umi-lidar-sdk-e2eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
1200,
1920,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/robot-umi-lidar-sdk-e2e.take_lid_off_saucepan_25_08_21_lerobotv2.1dual-lidar-umi-relative-filtered
Filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi-relative. It contains 65 demonstrations (63383 frames) that pass the complete continuous YAM replayability classification.
observation.state retains the original 12-D UMI Cartesian schema. Its values are the smoothed, trajectory-optimized YAM-achievable FK path mapped back into the UMI coordinate convention. Gripper observations, videos, timestamps, and tasks are preserved… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-relative-filtered.opus_lid_filtered
What this dataset is
This dataset is a filtered version of the MaLA-LM/mala-opus-dedup-2410 dataset after having all texts with incorrect language codes filtered out of it.
To do this, we use the cis-lmu/glotlid language ID model, predict the top probability language for both texts (source and target) and discard a row in which either text has a different language code to its original label.
How this dataset was made
from tqdm.auto import tqdm
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/opus_lid_filtered.dual-lidar-umi-independentThis dataset was created using LeRobot.
Dataset Description
Standalone LeRobot v3 dataset containing 296 dual-UMI orange-collection demonstrations (276,332 frames, 2.559 hours at 30 FPS). It contains synchronized observation.images.umi1 and observation.images.umi2 video observations, 14D observation.state, and 14D action; no LiDAR files or LiDAR frame features are included.
For each UMI independently, pose is expressed relative to that UMI's episode-start pose using… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-independent.dual-lidar-umi-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-test.tartanground_lidarremove-pen-lid-2_20260821_135420This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"right_joint_1.pos",
"right_joint_2.pos",
"right_joint_3.pos",
"right_joint_4.pos",
"right_joint_5.pos",
"right_joint_6.pos",
"right_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/reece-omahoney/remove-pen-lid-2_20260821_135420.pick_up_lid_from_pot_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 53,
"total_frames": 47362,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Grimster/pick_up_lid_from_pot_v2.close-lid-successThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"delta_x",
"delta_y",
"delta_z",
"delta_rx",
"delta_ry",
"delta_rz",
"gripper.pos"
],
"shape": [
7… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-vislab/close-lid-success.mala-opus-dedup-2410-lid-filteredclose-lid-failThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"delta_x",
"delta_y",
"delta_z",
"delta_rx",
"delta_ry",
"delta_rz",
"gripper.pos"
],
"shape": [
7… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-vislab/close-lid-fail.earbuds_case_assembly_with_lid_operationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 40,
"total_frames": 162924,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Vertax/earbuds_case_assembly_with_lid_operation.
