datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NExTQALLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.NExTQALinear-Next-Datasets
Linear Next Benchmark
Linear Next is a comprehensive benchmark designed to fairly compare various efficient transformer architectures. This project evaluates different approaches including linear attention, sparse attention, and other model structures under identical training conditions and datasets.
Overview
The benchmark aims to provide an unbiased comparison of efficient transformer variants by ensuring all models are trained with the same datasets, hyperparameters… See the full description on the dataset page: https://huggingface.co/datasets/Linear-Next/Linear-Next-Datasets.za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.Next_Token_Prediction_datasetgenai-image-tag-db
GenAI Image Tag DB (cc0-1.0)
This repository contains the cc0-1.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc0.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db.LLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/AlayaNeW/LLaVA-NeXT-Data.libero-pickandplace-segment-next-scene-ab-2genai-image-tag-db-CC4
GenAI Image Tag DB (cc-by-4.0)
This repository contains the cc-by-4.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc4.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-CC4.fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.genai-image-tag-db-mit
GenAI Image Tag DB (mit)
This repository contains the mit build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-mit.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects and health… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-mit.LLaVA-NeXT-Data-Reformattedllava-next-data-400k
Source & citation
This subset is derived from lmms-lab/LLaVA-NeXT-Data.
@misc{liu2024llavanext,
title={LLaVA-NeXT: Improved reasoning, OCR, and world knowledge},
url={https://llava-vl.github.io/blog/2024-01-30-llava-next/},
author={Liu, Haotian and Li, Chunyuan and Li, Yuheng and Li, Bo and Zhang, Yuanhan and Shen, Sheng and Lee, Yong Jae},
month={January},
year={2024}
}
Next.js-DatasetGR1-Tabletop-NextState-100x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-100x24.task1729_personachat_generate_next
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1729_personachat_generate_next
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1729_personachat_generate_next.Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.Qwen3-Coder-Next-Open-Code-SFT
Dataset Card — OpenCode Rejection Sampling
Overview
This dataset contains high-quality code reasoning data for training language models on competitive programming tasks. It is produced via rejection sampling with Qwen3-Coder-Next, which would generate multiple candidate solutions per problem, each candidate is executed against test cases in a sandboxed environment, and the results are used to build two complementary training datasets:
SFT dataset (49,374 examples)… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-Open-Code-SFT.Medical-Reasoning-SFT-Qwen3-Next-80B
Medical-Reasoning-SFT-Qwen3-Next-80B
A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
Qwen/Qwen3-Next-80B-A3B-Thinking
Total Samples
604,249
Samples with Reasoning
604,249 (100%)
Estimated Tokens
~1.42 Billion
Content Tokens
~505 Million
Reasoning Tokens
~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.depth-model-training-pack-next-pack-497ced43-6926d7be
Indoor Laptop Detection
Training dataset of laptops rendered across indoor home and office scenes, with RGB, albedo and metric depth plus per-frame annotations, built to train a laptop object-detection model.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/depth-model-training-pack-next-pack-497ced43-6926d7be.terminal-wm-sft-nextobs-v2
Terminal World-Model SFT (nextobs, corrected v2)
Corrected supervised fine-tuning data derived from
open-thoughts/OpenThoughts-Agent-SFT-100K
Terminus traces. Schema version: v2-observation-action.
Causal format
Every world-model transition is serialized as:
system: target-specific world-model instruction
user: task context (turn 1) + real current observation_t + executed action_t
assistant: target derived from the real observation_t+1
Rows remain multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/terminal-wm-sft-nextobs-v2.llava-next-data-100klighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230
Home Object Detection, Grasping and Sorting Eval — YOLOv8
Evaluation dataset for a robotic arm that detects, grasps, and sorts objects by type in home environments. 30 renders at 640x640 across kitchen, entry, living room and dressing spaces, staged with everyday household objects. Includes RGB plus metric depth, world-space normals (OpenGL, linear), albedo and material index passes, per-frame annotations, and midday lighting. Targets a YOLOv8 model.
This dataset mirrors public… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230.door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805
Outdoor Door Handle Push/Pull Training Set
Synthetic outdoor dataset staged in an alleyway to train a robot to detect door handles and determine whether a door must be pushed or pulled. 20 renders at 1024x1024 with albedo and metric depth passes, per-frame annotations, and authored lighting.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805.global-job-postings-multi-ats
Multi-ATS Job Postings Snapshot (June 2026)
A dated snapshot of 112,816 job postings collected from 23 different
Applicant Tracking Systems (Workday, Greenhouse, Lever, Ashby, iCIMS,
SmartRecruiters, Oracle Cloud, Workable, Paylocity, and more) and parsed into a
clean, consistent 47-column schema with a large language model.
Unlike typical single-source job dumps, every row is enriched with extracted
skills, salary ranges, qualifications, experience level, work model, and… See the full description on the dataset page: https://huggingface.co/datasets/NextGig-Rocks/global-job-postings-multi-ats.generations-qwen3-coder-next-pre_vallibero_spatial_next_object_target_dot_alpha04This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/libero_spatial_next_object_target_dot_alpha04.construction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779
Aerial Outdoor Scenes for Drone Ice-on-Power-Line Detection (YOLO)
Training dataset for a drone-mounted YOLO model that detects ice accretion on power lines during aerial inspection. Renders are staged in outdoor open-air scenes at 1280x1280 with RGB, albedo and world-space normal (OpenGL, linear) passes, per-frame annotations (bounding boxes, camera pose, labels), and authored daylight. Note: the available environments do not frame actual power lines, so the set stands in for… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/construction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779.
