datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
easy_5000_data_processingmedium_5000_data_processingmedium_5000-data_processing_n100k1medium_5000_data_processing_fixednemotron-terminal-data_processing
nemotron-terminal-data_processing
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_processing". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_processing.easy_5000-data_processing_n100k1terminal_bench_2_nemotron_terminal_data_processing__Qwen3_8B_20260413_170737data_processing_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dual_arm_robot",
"total_episodes": 1,
"total_frames": 3105,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/data_processing_test.easy_5000_data_processing_fixedmixed_1000_data_processingmixed_1000-data_processing_1000_n30k1mixed_1000_data_processing_fixedswebench_verified_random_100_folders_nemotron_terminal_data_processing__Qwen3_87dd0272edev_set_v2_nemotron_terminal_data_processing__Qwen3_8B_20260413_175806
