datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.NexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.LLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.NExTQANexus-Agents-ToolCalling
Nexus Agents — Tool-Calling Conversations
Synthetic, schema-verified tool-calling conversations for training the Nexus Projects
agents. This is the exact data behind
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF),
including the verification transcripts that scored it (27/27 on the behavioral
interview eval, vs 13/27 for the base model).
Links: the fine-tuned model →
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) ·
the generator + seed data + eval harness →
Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.NExTQANExtLong-512K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-512K-dataset.solana-dex-dataNexus_Ulawengza-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.nexus-jobsNEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.NExtLong-64K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-64K-dataset.NExtLong-128K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-128K-dataset.Linear-Next-Datasets
Linear Next Benchmark
Linear Next is a comprehensive benchmark designed to fairly compare various efficient transformer architectures. This project evaluates different approaches including linear attention, sparse attention, and other model structures under identical training conditions and datasets.
Overview
The benchmark aims to provide an unbiased comparison of efficient transformer variants by ensuring all models are trained with the same datasets, hyperparameters… See the full description on the dataset page: https://huggingface.co/datasets/Linear-Next/Linear-Next-Datasets.SWE-Next
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
SWE-Next Dataset
SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-Next.Next_Token_Prediction_datasetgenai-image-tag-db
GenAI Image Tag DB (cc0-1.0)
This repository contains the cc0-1.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc0.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db.LLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/AlayaNeW/LLaVA-NeXT-Data.swedish-legal-decisions-raw-v1
Swedish Court Decisions — Svenska Domstolsavgöranden
55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training.
The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations.
Why This Dataset
Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.genai-image-tag-db-CC4
GenAI Image Tag DB (cc-by-4.0)
This repository contains the cc-by-4.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc4.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-CC4.fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.LLaVA-NeXT-Data-Reformattedgenai-image-tag-db-mit
GenAI Image Tag DB (mit)
This repository contains the mit build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-mit.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects and health… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-mit.llava-next-data-400k
Source & citation
This subset is derived from lmms-lab/LLaVA-NeXT-Data.
@misc{liu2024llavanext,
title={LLaVA-NeXT: Improved reasoning, OCR, and world knowledge},
url={https://llava-vl.github.io/blog/2024-01-30-llava-next/},
author={Liu, Haotian and Li, Chunyuan and Li, Yuheng and Li, Bo and Zhang, Yuanhan and Shen, Sheng and Lee, Yong Jae},
month={January},
year={2024}
}
Nexora-music-pd-v1-mediumQwen3.8-Flash-Next-GGUF-metrics
