datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-V4.1-Flash-NVFP4-metrics
DeepSeek-V4.1-Flash-NVFP4 metrics
Everything behind the numbers in AtomicChat/DeepSeek-V4.1-Flash-NVFP4-nvidia.
logprobs/lp-<run>-<corpus>.npz: the raw top-512 log probabilities of every measurement run, 49,152 scored
positions each: ref, ref-repeat, ref-r3, ref-b1 (batch size 1) for the original; flat, flat-r2,
flat-r3 for the uncalibrated cast; nvidia, nvidia-r2, nvidia-r3 for the calibrated checkpoint.
logs/kld-<run>-<corpus>.json: the KL lower bound per run against ref… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/DeepSeek-V4.1-Flash-NVFP4-metrics.put_pen_into_bag_V4.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 101,
"total_frames": 104050,
"total_tasks": 1,
"total_videos": 303,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/put_pen_into_bag_V4.1.plug-usb-v4.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 99,
"total_frames": 104774,
"total_tasks": 1,
"total_videos": 297,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/plug-usb-v4.1.put_plate_to_rack_v4.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 91,
"total_frames": 73328,
"total_tasks": 1,
"total_videos": 273,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:91"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/put_plate_to_rack_v4.1.Deepseek-v4.1-CoTDatasets created by deepseek-ai/DeepSeek-V4.1-Flash.
You can use them to distill other models.
Ty!
Stuffed_Animal_V4.1_3cam_Normal_bboxes
Stuffed_Animal_V4.1_3cam_Normal
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Stuffed_Animal_V4.1_3cam_Normal
Stuffed_Animal_V4.1_3cam_Normal
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
put_plate_to_rack_v4.1_low_resThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 91,
"total_frames": 73328,
"total_tasks": 1,
"total_videos": 273,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:91"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/put_plate_to_rack_v4.1_low_res.fold_v4.1
fold_v4.1
Task: fold_tshirt_side_no_drag_edge
Episodes: 200
Frames: 138733
FPS: 25
Collection: 8 workers × 25 episodes, with keep_failed=true.
YOLO2431_put_pen_into_bag_V4.1
Put Pen Into Bag V4.1 TsFile
This dataset is an Apache TsFile conversion of YOLO2431/put_pen_into_bag_V4.1,
a LeRobot v2.1 yam_bimanual robot-manipulation dataset.
Source and provenance
Original dataset: YOLO2431/put_pen_into_bag_V4.1
Inspected source revision: 74df55d31492ff524fab30f0e2812d052dc4afcc
Original author/uploader: Sam F (YOLO2431)
License: Apache-2.0
Task (task_index = 0): Unzip the bag, pick up the pens from the table one at a time and place them… See the full description on the dataset page: https://huggingface.co/datasets/THULab/YOLO2431_put_pen_into_bag_V4.1.eval_act_so100_v4.1_1c_v1.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 8935,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seonixx/eval_act_so100_v4.1_1c_v1.1.mathlib_informal_v4.16.0
Notes
Names
All names in Lean (names of symbols and modules) are stored as their raw form (list[int | str]) instead of the usual pretty-printed form to avoid problems arising from quoting/unquoting.
For example, instead of "Lean.«binderTerm∉_»" we have ["Lean", "binderTerm∉_"].
Astraea_Chat_v4.1
🧭 Astraea Chat Dataset — v4.1
Synthetic Prompt-Architect Conversations for Meta-Prompting Fine-Tunes
The training data behind the astraea-chat-v10 prompt-architect model.
⚠️ Status & Disclaimer
This dataset is a work in progress and is not the final release. v4.1 is an intermediate snapshot used for active development and fine-tuning experiments. Schema, label distribution, formatting conventions, and record count are all subject to… See the full description on the dataset page: https://huggingface.co/datasets/braydenh563/Astraea_Chat_v4.1.details_bardsai__jaskier-7b-dpo-v4.1
Dataset Card for Evaluation run of bardsai/jaskier-7b-dpo-v4.1
Dataset automatically created during the evaluation run of model bardsai/jaskier-7b-dpo-v4.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bardsai__jaskier-7b-dpo-v4.1.eval_act_so100_v4.1_1cThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 17852,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seonixx/eval_act_so100_v4.1_1c.act_so100_v4.1_1cThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 22500,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seonixx/act_so100_v4.1_1c.so101_2cam_red_cube_v4.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 3,
"total_frames": 1309,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/L7-Robotics/so101_2cam_red_cube_v4.1.YOLO2431_put_pen_into_bag_V4.1
Put Pen Into Bag V4.1 TsFile
This dataset is an Apache TsFile conversion of YOLO2431/put_pen_into_bag_V4.1,
a LeRobot v2.1 yam_bimanual robot-manipulation dataset.
Source and provenance
Original dataset: YOLO2431/put_pen_into_bag_V4.1
Inspected source revision: 74df55d31492ff524fab30f0e2812d052dc4afcc
Original author/uploader: Sam F (YOLO2431)
License: Apache-2.0
Task (task_index = 0): Unzip the bag, pick up the pens from the table one at a time and place them… See the full description on the dataset page: https://huggingface.co/datasets/azithromycin/YOLO2431_put_pen_into_bag_V4.1.semiconductor_ins_v4.1_filtered_formatedmathlib_informal_v4.15.0
mathlib_informal_v4.15.0
Dataset Summary
This dataset contains Lean v4.15.0 mathlib declarations with informal descriptions produced by the Autoprover enrichment pipeline and published in the retrieval schema used by this codebase.
What Is Included
mathlib_informal_v4.15.0.jsonl: one JSON object per declaration
dataset_metadata.json: supplemental provenance, schema, and checksum metadata
Cleaning And Normalization
Machine-local paths were removed… See the full description on the dataset page: https://huggingface.co/datasets/adeo1/mathlib_informal_v4.15.0.Stuffed_Animal_V4.1_3cam_Fallen
Stuffed_Animal_V4.1_3cam_Fallen
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
mew1a-v4.1-pokemon-tcg-ultimate-with-redditmathlib_informal_v4.19.0func_calling_v4.1Granite-v4.1-Distilled-15K
⛰️ Granite-v4.1-Distilled-15K
Dataset Summary
Granite-v4.1-Distilled-15k is a supervised fine-tuning dataset for logic-oriented distillation. The prompts to the questions come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated using the only granite-4.1-8b teaching model.
Dataset Details
Dataset
constructai/Granite-v4.1-Distilled-15K
Source questions
Jackrong/GLM-5.1-Reasoning-1M-Cleaned
Teacher model
Granite-4.1-8b… See the full description on the dataset page: https://huggingface.co/datasets/constructai/Granite-v4.1-Distilled-15K.Stuffed_Animal_V4.1_3cam_Stuck
Stuffed_Animal_V4.1_3cam_Stuck
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
v4.1-nightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 10928,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/skpro19/v4.1-night.ssf-dataset-synthetic-v4.1
Dataset Card for ssf-dataset-synthetic-v4.1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dnth/ssf-dataset-synthetic-v4.1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dnth/ssf-dataset-synthetic-v4.1.cwec-v4.14-weaknesses-1.0
Introduction
This dataset is based on the complete XML file of CWE List Version 4.14 and is intended to provide researchers and security experts with structured data on Common Weakness Enumeration (CWE) for software and hardware. The dataset contains 963 entries in Alpaca format, each providing detailed information about a specific weakness.
Dataset Structure
Each entry in the dataset includes the following fields:
ID: The unique identifier for the weakness (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bayuncao/cwec-v4.14-weaknesses-1.0.Stuffed_Animal_V4.1_3cam_Merge2NFS_bboxes
Stuffed_Animal_V4.1_3cam_Merge2NFS
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
