datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GR1-Tabletop-NextState-1000x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-1000x24.LLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.NExTQAfglxdgLLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.ffNExTQAdsNExT-GQA
Can I Trust Your Answer? Visually Grounded Video Question Answering
Introduction
We study visually grounded VideoQA by forcing vision-language models (VLMs) to answer questions and simultaneously ground the relevant video moments as visual evidences. We show that this task is easy for human yet is extremely challenging for existing VLMs, revealing that the strong QA performance of these models may largely due to short-cut learning (e.g., language priors and spurious vision-text… See the full description on the dataset page: https://huggingface.co/datasets/jinyoungkim/NExT-GQA.TAT-QA
TAT-QA
Project Page
Paper - ACL 21
Paper - Arxiv
Source Code
Leaderboard
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
The unique features of TAT-QA include:
The context given is hybrid, comprising a semi-structured table and at least two relevant paragraphs that describe, analyze or… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/TAT-QA.BLIP3o-NEXT-EDIT-ENSEMBLE-DATASETSnextgqa
NExT-GQA (mirror)
A redistribution of the NExT-GQA benchmark, packaged as a single self-contained
repo (annotations + the 1,570 videos needed to run it) for convenience.
This is not the official release. All credit goes to the original authors.
Official code and data: https://github.com/doc-doc/NExT-GQA
Dataset description
NExT-GQA extends NExT-QA with temporal
grounding labels: for each multiple-choice question it annotates the video
segment(s) that actually… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/nextgqa.nextqa-rawvideoNeXTVideoA copy of the NextQA dataset, used to demonstrate fine-tuning Aria on video dataset.
pred_llava_next_10kNExtLong-512K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-512K-dataset.za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.LLaVA-NeXT-780k-webdatasetStage-2: Visual Instruction Tuning, this is a subset for quick start.
nextturn-study-clipsNExtLong-64K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-64K-dataset.DA-Next-5M-exampleNExTVideoNExtLong-128K-dataset
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-128K-dataset.Linear-Next-Datasets
Linear Next Benchmark
Linear Next is a comprehensive benchmark designed to fairly compare various efficient transformer architectures. This project evaluates different approaches including linear attention, sparse attention, and other model structures under identical training conditions and datasets.
Overview
The benchmark aims to provide an unbiased comparison of efficient transformer variants by ensuring all models are trained with the same datasets, hyperparameters… See the full description on the dataset page: https://huggingface.co/datasets/Linear-Next/Linear-Next-Datasets.SWE-Next
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
SWE-Next Dataset
SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-Next.TAT-DQA
TAT-DQA
Project Page
Paper - MM 22
Paper - Arxiv
Github
Leaderboard
TAT-DQA is a large-scale Document VQA dataset, which is constructed by extending the TAT-QA. It aims to stimulate the progress of QA research over more complex and realistic visually-rich documents with rich tabular and textual content, especially those requiring numerical reasoning.
The unique features of TAT-DQA include:
The documents in TAT-DQA dataset are sampled from real-world high-quality financial… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/TAT-DQA.Next_Token_Prediction_datasetqwen3.8-flash-next-expert-traces
Qwen3.8-Flash-Next expert routing traces
Token-level routing traces of a deployed MoE model: for every token and every one of the
48 MoE layers, which experts the router chose, the top-32 router logits behind that choice,
and the exact hidden state the router read — plus, in v3, the state at many layers per token,
the post-final-norm state the LM head consumes, and the LM head's top-8 next-token candidates.
The corpus exists to answer one question: how well can the next tokens'… See the full description on the dataset page: https://huggingface.co/datasets/aswinkumar99/qwen3.8-flash-next-expert-traces.
