datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GR1-Tabletop-NextState-1000x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-1000x24.NExTQANExTQAgenai-image-tag-db
GenAI Image Tag DB (cc0-1.0)
This repository contains the cc0-1.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc0.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db.genai-image-tag-db-CC4
GenAI Image Tag DB (cc-by-4.0)
This repository contains the cc-by-4.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc4.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-CC4.fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.genai-image-tag-db-mit
GenAI Image Tag DB (mit)
This repository contains the mit build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-mit.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects and health… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-mit.NextNightQwen3.8-Flash-Next-GGUF-metricsGR1-Tabletop-NextState-100x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders, 185… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-100x24.glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root.
first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds
the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files
are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset.
GLM5-Next tiny native CPU fixture
This is a complete untrained random-initialized native Glm5NextForConditionalGeneration
wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.Qwen3-Coder-Next-Open-Code-SFT
Dataset Card — OpenCode Rejection Sampling
Overview
This dataset contains high-quality code reasoning data for training language models on competitive programming tasks. It is produced via rejection sampling with Qwen3-Coder-Next, which would generate multiple candidate solutions per problem, each candidate is executed against test cases in a sandboxed environment, and the results are used to build two complementary training datasets:
SFT dataset (49,374 examples)… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-Open-Code-SFT.depth-model-training-pack-next-pack-497ced43-6926d7be
Indoor Laptop Detection
Training dataset of laptops rendered across indoor home and office scenes, with RGB, albedo and metric depth plus per-frame annotations, built to train a laptop object-detection model.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/depth-model-training-pack-next-pack-497ced43-6926d7be.global-job-postings-multi-ats
Multi-ATS Job Postings Snapshot (June 2026)
A dated snapshot of 112,816 job postings collected from 23 different
Applicant Tracking Systems (Workday, Greenhouse, Lever, Ashby, iCIMS,
SmartRecruiters, Oracle Cloud, Workable, Paylocity, and more) and parsed into a
clean, consistent 47-column schema with a large language model.
Unlike typical single-source job dumps, every row is enriched with extracted
skills, salary ranges, qualifications, experience level, work model, and… See the full description on the dataset page: https://huggingface.co/datasets/NextGig-Rocks/global-job-postings-multi-ats.lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230
Home Object Detection, Grasping and Sorting Eval — YOLOv8
Evaluation dataset for a robotic arm that detects, grasps, and sorts objects by type in home environments. 30 renders at 640x640 across kitchen, entry, living room and dressing spaces, staged with everyday household objects. Includes RGB plus metric depth, world-space normals (OpenGL, linear), albedo and material index passes, per-frame annotations, and midday lighting. Targets a YOLOv8 model.
This dataset mirrors public… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230.occluded-vegetables-in-packed-fridge-rt-detr-detection-next-pack-e522b0f8-7563b26f
Occluded Vegetable Detection in Packed Fridge
Training dataset of packed refrigerator scenes with partially occluded vegetables, rendered to fine-tune an RT-DETR object detection model for improved detection under heavy occlusion.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/occluded-vegetables-in-packed-fridge-rt-detr-detection-next-pack-e522b0f8-7563b26f.door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805
Outdoor Door Handle Push/Pull Training Set
Synthetic outdoor dataset staged in an alleyway to train a robot to detect door handles and determine whether a door must be pushed or pulled. 20 renders at 1024x1024 with albedo and metric depth passes, per-frame annotations, and authored lighting.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805.nextplase_projectgenerations-qwen3-coder-next-pre_valconstruction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779
Aerial Outdoor Scenes for Drone Ice-on-Power-Line Detection (YOLO)
Training dataset for a drone-mounted YOLO model that detects ice accretion on power lines during aerial inspection. Renders are staged in outdoor open-air scenes at 1280x1280 with RGB, albedo and world-space normal (OpenGL, linear) passes, per-frame annotations (bounding boxes, camera pose, labels), and authored daylight. Note: the available environments do not frame actual power lines, so the set stands in for… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/construction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779.outdoor-door-handle-push-pull-training-set-next-pack-1c340bea-fea6c091
Door Handle Detection & Push/Pull Direction — Training Pack
Training dataset to teach a robot to locate door handles and infer whether each door must be pushed or pulled to open. Renders staged in home and office environments (office, home entrance, break area) that frame doors with various handle types. Outputs RGB with metric depth and world-space normals for contact geometry, per-frame annotations (bounding boxes, camera pose), authored lighting, at 1024×1024. Targets object… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/outdoor-door-handle-push-pull-training-set-next-pack-1c340bea-fea6c091.kitchen-object-grasping-manipulation-policy-training-next-pack-5534d90d-f7cb0d04
Cluttered Home Kitchens for Cooking-Assistant Robot
Training dataset of richly cluttered, realistic home kitchens to train a YOLOv8-based cooking-assistant robot that navigates the kitchen, detects and segments objects, distinguishes food from non-food, identifies what needs cleaning, avoids obstacles and manipulates objects. Renders are 640x640 with metric depth, world-space normals (OpenGL linear), albedo and material-index passes, per-frame annotations, and midday lighting… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/kitchen-object-grasping-manipulation-policy-training-next-pack-5534d90d-f7cb0d04.fruit-on-table-recognition-for-sorting-next-pack-9a889d4d-aee688d0
Outdoor Fruit Detection & Sorting
Training dataset of 100 outdoor scenes (market squares and an alleyway) framing fruit among stalls and produce, at 640x640 with albedo, material index, world normals and metric depth passes plus per-frame annotations, to train a YOLOv8 model to detect, classify and count fruit types for eventual sorting.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fruit-on-table-recognition-for-sorting-next-pack-9a889d4d-aee688d0.glm5-next-tiny-fidelity-root-v1
glm5_next random CPU fixture root
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/glm5-next-tiny-random-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-fidelity-root-v1.turkish-plu-next-event-predictionHomepage: https://github.com/GGLAB-KU/turkish-plu
indoor-laptop-detection-next-pack-e6bd6a07-356e7b31
Indoor Plant Detection for YOLO
Synthetic training dataset of 5 renders (640x640) staged in home and office interiors — libraries and residential lounges — each framing an indoor plant. Includes RGB, albedo and metric depth passes with per-frame bounding-box annotations, for training a YOLO object-detection model to detect indoor plants.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/indoor-laptop-detection-next-pack-e6bd6a07-356e7b31.Qwen3-Coder-Next-Open-Code-SFT
Dataset Card — OpenCode Rejection Sampling
Overview
This dataset contains high-quality code reasoning data for training language models on competitive programming tasks. It is produced via rejection sampling with Qwen3-Coder-Next, which would generate multiple candidate solutions per problem, each candidate is executed against test cases in a sandboxed environment, and the results are used to build two complementary training datasets:
SFT dataset (49,374 examples)… See the full description on the dataset page: https://huggingface.co/datasets/derekib/Qwen3-Coder-Next-Open-Code-SFT.MVEmo
MVEmo
This is the dataset repository for the paper: Bridging Categorical and Dimensional Affect: The MVEmo Multi-Task Benchmark for Music-Related Emotion Recognition.
Dataset Details
Dataset Description
MVEmo is a large-scale multimodal dataset that consists of 11,764 music video samples with both static and dynamic emotion annotations for music-related emotion recognition (MRER). It consists of the following key features:
Basic Information: title, artist… See the full description on the dataset page: https://huggingface.co/datasets/NEXTLab-ZJU/MVEmo.llm-agent-harness-reliability-next-prime
LLM Next Prime Harness Dataset
This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool.
The task is deliberately simple and objectively verifiable:
What is the smallest prime number that is strictly greater than n?
The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.
