datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.Medical-Reasoning-SFT-Qwen3-Next-80B
Medical-Reasoning-SFT-Qwen3-Next-80B
A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
Qwen/Qwen3-Next-80B-A3B-Thinking
Total Samples
604,249
Samples with Reasoning
604,249 (100%)
Estimated Tokens
~1.42 Billion
Content Tokens
~505 Million
Reasoning Tokens
~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.tat-llm-instructions
TAT-LLM-Instructions
The TAT(Tabular and Textual)-LLM-Instructions dataset is a curated collection of financial data, structured to resemble instructions. It aggregates information from three publicly available tabular and textual QA datasets: FinQA, TAT-QA, and TAT-DQA. By employing specialized templates, TAT-LLM-Instructions transforms the original dataset into prompts that are optimized for compatibility with large language models (LLMs) and external executor, aiming to… See the full description on the dataset page: https://huggingface.co/datasets/next-tat/tat-llm-instructions.next.js-15.4-with-reasoning
Description
The Next.js Documentation Dataset based on next.js 15.4 version is a high-quality, code-centric dataset created from Next.js documentation for fine-tuning language models. It contains 1,172 question-answer pairs derived from 178 markdown documentation files, focusing on practical code examples and real-world development scenarios.
This dataset is designed for:
Question Answering: Natural language questions about Next.js development
Code Generation: Generating practical… See the full description on the dataset page: https://huggingface.co/datasets/Slava32/next.js-15.4-with-reasoning.omnimcp_nextjs_react_architect_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.NextSearch-1-Tasks
NextSearch-1 Tasks
The task pools behind the NextSearch-1
web research agents: every row is a research question with its reference
answer and grading spec — the sft-tasks configs are the tasks behind the
supervised corpora, the rl-tasks configs the verified prompt+gold pools
used for reinforcement learning. Full trajectories for the SFT configs are in the companion
NextSearch-1-Trajectories.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.NextGenBench
Dataset Card for Next Generation Benchmark
Dataset Summary
This is a multitask test consisting of only questions (some are MCQ) from various branches of knowledge. Specifically the following topics:
Abstract Algebra
Anatomy
Astronomy
Business Ethics
Clinical Knowledge
Primary School Biology
Primary School Chemistry
Primary School Physics
Primary School Math
Primary School English
Primary School Science
Primary School Computer Science
Computer Security
Daily Economy… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/NextGenBench.docs-instruct-nextjs-20260601-0306
docs-instruct-20260601-0306
Synthetic instruction-tuning dataset generated by the DownFTuner pipeline.
Source: random Wikipedia articles (en), one run.
Generator: LLM-synthesized instruction/answer pairs grounded in each article.
Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl.
License: CC-BY-SA-4.0 (inherits from Wikipedia source).
Source URLs are preserved in each row's source metadata.
nextjs-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
nextjs-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
NextGenAI
