CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /SWE-Next SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next Dataset SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-Next.texttext-generation1K<n<10K1 likes685 downloads5mo agoHugging Face02TIGER-Lab /SWE-Next-SFT-Trajectories SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next SFT Trajectories SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests. The dataset is designed for training… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-Next-SFT-Trajectories.texttext-generation1K<n<10K3 likes263 downloads5mo agoHugging Face03NextTokenAI /NextSearch-1-Trajectories NextSearch-1 Trajectories The supervised training trajectories behind the NextSearch-1 web research agents: complete research episodes — reasoning, tool calls, live-web tool results, and final answers — for every task in the companion NextSearch-1-Tasks SFT configs. Directly trainable: each row is a prompt (messages) plus a target trajectory (target) with per-message reasoning and OpenAI-format tool calls. Technical report: nexttoken.co/research/nextsearch-1 · Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Trajectories.texttext-generation1K<n<10K0 likes135 downloads1mo agoHugging Face04Slava32 /next.js-15.4-with-reasoning Description The Next.js Documentation Dataset based on next.js 15.4 version is a high-quality, code-centric dataset created from Next.js documentation for fine-tuning language models. It contains 1,172 question-answer pairs derived from 178 markdown documentation files, focusing on practical code examples and real-world development scenarios. This dataset is designed for: Question Answering: Natural language questions about Next.js development Code Generation: Generating practical… See the full description on the dataset page: https://huggingface.co/datasets/Slava32/next.js-15.4-with-reasoning.textquestion-answering1K<n<10K1 likes93 downloads1y agoHugging Face05NextGenInstitute /socraticDataset1680 Socratic AI Pedagogy Preference Dataset (1,680 Quadruplets) A curated multi-domain educational preference dataset for post-training open language models into pedagogical Socratic tutors for introductory Artificial Intelligence and Machine Learning courses. 📚 Dataset Overview The dataset contains 1,680 paired preference quadruplets across five foundational AI subfields: Classical Search & Planning (347 items): A* heuristic admissibility, graph search state-space… See the full description on the dataset page: https://huggingface.co/datasets/NextGenInstitute/socraticDataset1680.texttext-generation1K<n<10K0 likes87 downloads1mo agoHugging Face06NextTokenAI /NextSearch-1-Tasks NextSearch-1 Tasks The task pools behind the NextSearch-1 web research agents: every row is a research question with its reference answer and grading spec — the sft-tasks configs are the tasks behind the supervised corpora, the rl-tasks configs the verified prompt+gold pools used for reinforcement learning. Full trajectories for the SFT configs are in the companion NextSearch-1-Trajectories. Technical report: nexttoken.co/research/nextsearch-1 · Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.textquestion-answering10K<n<100K2 likes77 downloads1mo agoHugging Face07NextGenC /synapse-set-10k 🧠 SynapseSet-10K SynapseSet-10K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-10k.texttext-generation10K<n<100K1 likes44 downloads1y agoHugging Face08NextGenC /synapse-set-50k 🧠 SynapseSet-50K SynapseSet-50K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-50k.texttext-generation10K<n<100K1 likes42 downloads1y agoHugging Face09NextGenC /synapse-set-100k 🧠 SynapseSet-100K SynapseSet-100K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-100k.texttext-generation100K<n<1M2 likes41 downloads1y agoHugging Face10Evelina2025 /SWE-Next SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next Dataset SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/Evelina2025/SWE-Next.texttext-generation1K<n<10K0 likes28 downloads5mo agoHugging Face11psychopenguin /next_token Supreme Court of India Judgments Dataset (1950-2025) Dataset Description This dataset contains a comprehensive collection of judgments and orders from the Supreme Court of India, spanning from its inception in 1950 up to early 2025. Dataset Summary Total Documents: 26,688 Total Tokens: ~196.9 Million (counted using cl100k_base encoding) Format: JSONL (JSON Lines) Language: English Time Range: 1950 - 2025 Data Fields Each entry in the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/psychopenguin/next_token.texttext-generation10K<n<100K0 likes22 downloads9mo agoHugging Face12alucent /mirror-SWE-Next-SFT-Trajectoriesgated SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next SFT Trajectories SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-SWE-Next-SFT-Trajectories.texttext-generation1K<n<10K0 likes21 downloads2mo agoHugging Face13emdadibos /nextjobzlm-dataset-v0.1.0 NextjobzLM Instruction Dataset v0.1.0 Training data for NextjobzLM, the explanation and extraction model behind NextJobz recommendations. This is the v0.1.0 release, published on its own so the dataset viewer and fine-tuning UIs can read it. Every row is generated by code in the repository; no real person's data is in here. The model this trains never scores, ranks, or decides eligibility. Those decisions come from a rule engine and a ranker; the model restates them and extracts… See the full description on the dataset page: https://huggingface.co/datasets/emdadibos/nextjobzlm-dataset-v0.1.0.texttext-generation10K<n<100K0 likes16 downloads2mo agoHugging Face14jokernifty /docs-instruct-nextjs-20260601-0306 docs-instruct-20260601-0306 Synthetic instruction-tuning dataset generated by the DownFTuner pipeline. Source: random Wikipedia articles (en), one run. Generator: LLM-synthesized instruction/answer pairs grounded in each article. Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl. License: CC-BY-SA-4.0 (inherits from Wikipedia source). Source URLs are preserved in each row's source metadata. texttext-generation1K<n<10K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.