datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nextjs-selfbench
Next.js Selfbench
Next.js Selfbench is a 47-task software-engineering evaluation packaged for the Harbor evaluation framework. Each task asks an agent to implement a change in a frozen revision of vercel/next.js, then evaluates the resulting patch with task-specific tests.
This repository contains the raw evaluation only. It does not include model outputs, scores, costs, or benchmark result artifacts.
Use
Download the raw task package with the Hugging Face CLI:
hf… See the full description on the dataset page: https://huggingface.co/datasets/dari-ai/nextjs-selfbench.Next.js-Datasetnext.js-15.4-with-reasoning
Description
The Next.js Documentation Dataset based on next.js 15.4 version is a high-quality, code-centric dataset created from Next.js documentation for fine-tuning language models. It contains 1,172 question-answer pairs derived from 178 markdown documentation files, focusing on practical code examples and real-world development scenarios.
This dataset is designed for:
Question Answering: Natural language questions about Next.js development
Code Generation: Generating practical… See the full description on the dataset page: https://huggingface.co/datasets/Slava32/next.js-15.4-with-reasoning.omnimcp_nextjs_react_architect_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.nextjs-14omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.nextjs-react-dataset
Nextjs-React-Dataset
Made with ❤️ using 🦥 Unsloth Studio
Next.js 16 with React 19 was generated with Unsloth Recipe Studio. It contains 2,100 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("marianbusoi/nextjs-react-dataset", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 2,100
📋 Columns: 3
✅ Completion: 21.0% (10,000 requested)
📋 Schema… See the full description on the dataset page: https://huggingface.co/datasets/marianbusoi/nextjs-react-dataset.Next-JS-Docsurls included in this dataset: https://nextjs.org/docs/
nextjs-dev-datasetThis is a fork from https://huggingface.co/datasets/furia0928/nextjs-dev-dataset
nextjs_typescript_fim_datasetnextjs16-react19-dataset
Nextjs16-React19-Dataset
Made with ❤️ using 🦥 Unsloth Studio
nextjs16+react19 was generated with Unsloth Recipe Studio. It contains 10,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("marianbusoi/nextjs16-react19-dataset", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 10,000
📋 Columns: 1
📋 Schema & Statistics
Column
Type
Column… See the full description on the dataset page: https://huggingface.co/datasets/marianbusoi/nextjs16-react19-dataset.nextjs1000nextjs-app-docsNextjs-app-docsthis is a huggingface duplicate for further processing
nextjs-chakra-ui-datasetNext.js-Dataset-Convertednextjs16nextjs-sharegptdocs-instruct-nextjs-20260601-0306
docs-instruct-20260601-0306
Synthetic instruction-tuning dataset generated by the DownFTuner pipeline.
Source: random Wikipedia articles (en), one run.
Generator: LLM-synthesized instruction/answer pairs grounded in each article.
Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl.
License: CC-BY-SA-4.0 (inherits from Wikipedia source).
Source URLs are preserved in each row's source metadata.
nextjs-dev-datasetnextjs-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
nextjs-pluginnextjs-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Next.js
Documentation Data Source Link: https://nextjs.org/docs
Data Source License: https://github.com/vercel/next.js/blob/canary/license.md
Data Source Authors: Vercel
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
