datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TAT-QA
TAT-QA
Project Page
Paper - ACL 21
Paper - Arxiv
Source Code
Leaderboard
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
The unique features of TAT-QA include:
The context given is hybrid, comprising a semi-structured table and at least two relevant paragraphs that describe, analyze or… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/TAT-QA.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.nextgqa
NExT-GQA (mirror)
A redistribution of the NExT-GQA benchmark, packaged as a single self-contained
repo (annotations + the 1,570 videos needed to run it) for convenience.
This is not the official release. All credit goes to the original authors.
Official code and data: https://github.com/doc-doc/NExT-GQA
Dataset description
NExT-GQA extends NExT-QA with temporal
grounding labels: for each multiple-choice question it annotates the video
segment(s) that actually… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/nextgqa.TAT-DQA
TAT-DQA
Project Page
Paper - MM 22
Paper - Arxiv
Github
Leaderboard
TAT-DQA is a large-scale Document VQA dataset, which is constructed by extending the TAT-QA. It aims to stimulate the progress of QA research over more complex and realistic visually-rich documents with rich tabular and textual content, especially those requiring numerical reasoning.
The unique features of TAT-DQA include:
The documents in TAT-DQA dataset are sampled from real-world high-quality financial… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/TAT-DQA.swedish-legal-decisions-raw-v1
Swedish Court Decisions — Svenska Domstolsavgöranden
55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training.
The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations.
Why This Dataset
Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.MMDocBench
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
MMDocBench is an open-sourced benchmark with various OCR-free document understanding tasks for evaluating fine-grained visual perception and reasoning abilities.
For more details, please refer to the project page: https://MMDocBench.github.io/.
Dataset Structure
MMDocBench consists of 15 main tasks and 48 sub-tasks, involving 2,400 document images, 4,338 QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/MMDocBench.supra-nexus-o1-training
Supra Nexus O1 Training Datasets
Overview
Comprehensive training datasets for Supra Nexus O1 models, including:
Identity training
Chain-of-thought reasoning
Self-improvement examples (O1.5)
Instruction following
Datasets Included
1. Identity Dataset (supra_identity.jsonl)
Model identity and alignment
Organization information
Capability descriptions
2. Instruction Dataset (supra_instruct_*.jsonl)
Direct instruction… See the full description on the dataset page: https://huggingface.co/datasets/Supra-Nexus/supra-nexus-o1-training.Medical-Reasoning-SFT-Qwen3-Next-80B
Medical-Reasoning-SFT-Qwen3-Next-80B
A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
Qwen/Qwen3-Next-80B-A3B-Thinking
Total Samples
604,249
Samples with Reasoning
604,249 (100%)
Estimated Tokens
~1.42 Billion
Content Tokens
~505 Million
Reasoning Tokens
~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.tat-llm-instructions
TAT-LLM-Instructions
The TAT(Tabular and Textual)-LLM-Instructions dataset is a curated collection of financial data, structured to resemble instructions. It aggregates information from three publicly available tabular and textual QA datasets: FinQA, TAT-QA, and TAT-DQA. By employing specialized templates, TAT-LLM-Instructions transforms the original dataset into prompts that are optimized for compatibility with large language models (LLMs) and external executor, aiming to… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/tat-llm-instructions.next.js-15.4-with-reasoning
Description
The Next.js Documentation Dataset based on next.js 15.4 version is a high-quality, code-centric dataset created from Next.js documentation for fine-tuning language models. It contains 1,172 question-answer pairs derived from 178 markdown documentation files, focusing on practical code examples and real-world development scenarios.
This dataset is designed for:
Question Answering: Natural language questions about Next.js development
Code Generation: Generating practical… See the full description on the dataset page: https://huggingface.co/datasets/Slava32/next.js-15.4-with-reasoning.omnimcp_nextjs_react_architect_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.NextSearch-1-Tasks
NextSearch-1 Tasks
The task pools behind the NextSearch-1
web research agents: every row is a research question with its reference
answer and grading spec — the sft-tasks configs are the tasks behind the
supervised corpora, the rl-tasks configs the verified prompt+gold pools
used for reinforcement learning. Full trajectories for the SFT configs are in the companion
NextSearch-1-Trajectories.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.nexttoken-model-1-dataset-sft
NextToken Model 1 SFT dataset (v4)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
846 scheme/product source documents across 57 schemes/products, chunked
into 1,445 passages.
v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.Open-LLaVA-NeXT-mix1M
Open-LLaVA-NeXT 1M Dataset Card
Dataset details
Dataset type: 1M SFT data for re-producing LLaVA-NeXT series.
We augmented the sharegpt4v_mix665k dataset with additional data. We have made every effort to align our training data with that of LLaVA-NeXT. However, we were unable to access the tens of thousands of real user interaction data that LLaVA-NeXT collected. As a result, we used 200K ALLaVA-Instruct-VFLAN-4V data as a substitute. Additionally, since TextVQA has been… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/Open-LLaVA-NeXT-mix1M.NextGenBench
Dataset Card for Next Generation Benchmark
Dataset Summary
This is a multitask test consisting of only questions (some are MCQ) from various branches of knowledge. Specifically the following topics:
Abstract Algebra
Anatomy
Astronomy
Business Ethics
Clinical Knowledge
Primary School Biology
Primary School Chemistry
Primary School Physics
Primary School Math
Primary School English
Primary School Science
Primary School Computer Science
Computer Security
Daily Economy… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/NextGenBench.Open-MedQA-Nexus
Open Nexus MedQA
This dataset combines various publicly available medical datasets like ChatDoctor, icliniq, etc., into a unified format for training and evaluating medical question-answering models.
Dataset Details
Open Nexus MedQA is a comprehensive dataset designed to facilitate the development of advanced medical question answering systems. It integrates diverse medical data sources, meticulously processed to provide a uniform format. The format includes:… See the full description on the dataset page: https://huggingface.co/datasets/exafluence/Open-MedQA-Nexus.nexa-science-multitask-balanced
Nexa Science Multitask Balanced
This dataset is a curated, instruction-formatted scientific multitask mixture for:
claim verification (<TASK:VERIFY>)
abstract-grounded biomedical QA (<TASK:QA>)
retrieval relevance re-ranking (<TASK:RERANK>)
Format
Each row is JSONL with:
{task, instruction, input, output, meta}
Splits Included
train_balanced_short.jsonl
val_balanced_short.jsonl
stats_balanced_short.json
Notes
QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.indonluThe IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.ilmihalsynaxarium
Ethiopian Synaxarium Dataset (የኢትዮጵያ ስንክሳር)
Dataset Description
This dataset contains the complete Ethiopian Orthodox Tewahedo Church Synaxarium (ስንክሳር) - a collection of daily readings commemorating saints, martyrs, and biblical events according to the Ethiopian calendar.
The Synaxarium is an essential liturgical book used daily in Ethiopian Orthodox churches, containing hagiographies and spiritual readings for each day of the year.
Key Features
Complete… See the full description on the dataset page: https://huggingface.co/datasets/Nexuss0781/synaxarium.islam-ilmihali-Omer-Nasuh-Bilmen
Ömer Nasuh Bilmen İslam İlmihali Dataseti
Açıklama
Bu dataset, Ömer Nasuh Bilmen'in İslam İlmihali kitabı kullanılarak oluşturulmuştur. Kitaptaki yazılar parçalara bölünmüş ve bu parçalara uygun sorular, NousResearch/Hermes-3-Llama-3.1-405B modeli kullanılarak otomatik olarak üretilmiştir. Bu dataset, İslam üzerine yapılacak araştırmalar, eğitim materyalleri veya çeşitli yapay zeka projeleri için faydalı olabilir. Dataset herkesin kullanımına açık olup, izin verilen… See the full description on the dataset page: https://huggingface.co/datasets/NexusV/islam-ilmihali-Omer-Nasuh-Bilmen.docs-instruct-nextjs-20260601-0306
docs-instruct-20260601-0306
Synthetic instruction-tuning dataset generated by the DownFTuner pipeline.
Source: random Wikipedia articles (en), one run.
Generator: LLM-synthesized instruction/answer pairs grounded in each article.
Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl.
License: CC-BY-SA-4.0 (inherits from Wikipedia source).
Source URLs are preserved in each row's source metadata.
nextjs-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
nextjs-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
NextGen_BotNextGenAI
