datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt_injection_cleaned_dataset
Dataset Card for "prompt_injection_cleaned_dataset"
More Information needed
my_dataset_cleaned
my_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
数据集信息
总episodes数: 47 (原48个,已移除episode_000000)
总帧数: 8,918
任务: 抓取立方体并放入盒子
机器人: so-100
帧率: 30 FPS
数据质量说明
注意: 原始数据集中的第0个episode (episode_000000) 由于视频质量问题已被移除。当前数据集从episode_000001开始,包含47个高质量的episode。… See the full description on the dataset page: https://huggingface.co/datasets/myzxyz/my_dataset_cleaned.cleaned_data
Cleaned Tabula Muris Senis Single-Cell Data and other aging datasets
This dataset contains LLM-cleaned single-cell transcriptomic annotations from the Tabula Muris Senis project, specifically for mouse tissues processed with SmartSeq2, and ALL OTHER DATASETS WITH AGING IN THE FILENAME :-) . The cleaning and annotation were performed using large language models (OpenAI and Claude), enabling enriched metadata and corrected cell type labels.
🧬 Over 1.3 million rows and 78.17 GB… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/cleaned_data.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.mbti-Personalitycafe-cleaned-datatool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified
Deep Research - Tulu SFT Data Cleaned Rectified
👥 Follow the Author
Supriti Vijay
Overview
This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.best_deal_02_cleaned_datamR3-Dataset-Cleanedso101-dataset-cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 79,
"total_frames": 34886,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:78"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/AndrewNoviello/so101-dataset-cleaned.lerobot_dsue_dataset_pollinate_3_v2_cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DSue/lerobot_dsue_dataset_pollinate_3_v2_cleaned.superkart-dataset-cleanedwhite_paper_ball_in_bin_merged_dataset_cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 28,
"total_frames": 20226,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/GautamR/white_paper_ball_in_bin_merged_dataset_cleaned.lekiwi-full-dataset-cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 150,
"total_frames": 64992,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CRPlab/lekiwi-full-dataset-cleaned.cleaned_hiring_dataset_qval_w_original_fixed_promptdrug_dataset_cleanedm1_preference_data_cleanedC4AI-preference-dataset-cleanedlerobot_dsue_dataset_pollinate_3_cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 315,
"total_frames": 141956,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:315"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DSue/lerobot_dsue_dataset_pollinate_3_cleaned.clustered_oss_dataset_cleanedcve-cwe-dataset-cleaned
CVE-CWE Dataset (Cleaned)
Cleaned version of the CVE-CWE dataset with only standard CWE classifications.
Dataset Source
Original Dataset: stasvinokur/cve-and-cwe-dataset-1999-2025
This dataset contains CVE (Common Vulnerabilities and Exposures) descriptions paired with their corresponding CWE (Common Weakness Enumeration) classifications from 1999-2025.
Cleaning Process
The original dataset contained 280,694 samples. We performed the following cleaning:… See the full description on the dataset page: https://huggingface.co/datasets/LorenzoNava/cve-cwe-dataset-cleaned.cleaned-datasetbaseball-stats-cleaned_venue_dataprompt_injection_cleaned_dataset
Dataset Card for "prompt_injection_cleaned_dataset"
More Information needed
amazon-review-dataset-cleanedmerged_output_qs_only_exact_dedup_90_cleaned_dataset_morethan10resp_clip16_sortedcleaned-dataset-QMsumprompt_injection_cleaned_dataset
Dataset Card for "prompt_injection_cleaned_dataset"
More Information needed
humaneval-datagen-run-1_best_att_50_sol_50_20250225_153517_cleanedhumaneval-datagen-run-4_best_att_50_sol_50_20250225_141616_cleanedfully-cleaned-telugu-english-dataset
