datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
procedural-pile
Procedural Pile: Procedural reasoning data SFT (and RL)
Procedural Pile is a synthetic corpus of verifiable reasoning problems generated by Reasoning Core. It is intended for continued pretraining, mid-training, and supervised fine-tuning.
Answers come from procedural generators and task-specific solvers or checkers, rather than language-model generation. The corpus spans mathematics, formal logic, planning, graphs, parsing, code, structured data, and other symbolic domains.… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-core/procedural-pile.DSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.task1390_wscfixed_coreference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1390_wscfixed_coreference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1390_wscfixed_coreference.task891_gap_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task891_gap_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task891_gap_coreference_resolution.symbolic-reasoning-env
⚠️DEPRECATED: PLEASE MOVE to hf.co/reasoning-core/procedural-pile
Reasoning Core ◉
Paper: Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
Code: GitHub Repository
reasoning-core is a text-based RLVR for LLM reasoning training.
It is centered on expressive symbolic tasks, including full fledged FOL, formal mathematics with TPTP, formal planning with novel domains, and syntax tasks.
Abstract
We introduce Reasoning Core, a new… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-core/symbolic-reasoning-env.task893_gap_fill_the_blank_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task893_gap_fill_the_blank_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task893_gap_fill_the_blank_coreference_resolution.opus-gpt-swe-frontier-core
SWE Base
Repository-level software engineering trajectories for training coding agents.
2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost
SWE-bench · debugging · patching · tools · agents
Overview
SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.ShareGPT-X
Dataset Summary
ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed.
Supported Tasks and Leaderboards
text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.COREX-18textCORE-18 Fulltext
Introducing the CORE-18 Full Text dataset, among the first well-maintained public datasets of CORE. CORE offers one of the largest collections of research papers, including supplementary metadata, to support Artificial Intelligence, Machine Learning research, and engineering projects. This dataset has gained significant attention among major corporations and research laboratories for Natural Language Processing research.
Recognizing the importance of accessibility… See the full description on the dataset page: https://huggingface.co/datasets/laion/COREX-18text.task892_gap_reverse_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task892_gap_reverse_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task892_gap_reverse_coreference_resolution.i-love-reading-pixiv-novels
Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels
Dataset Details
Dataset Description
This is a more or less raw dump of pixiv novel data (13,012,017 documents to be exact.)
Are you the hacker?
I scraped pixiv on the same day of the Kadokawa site issues. I had no clue about the issue surrounding nicolive, etc until I noticed after the scrape was done.
around 8 hours before I started the scrape, the websites(?) went down. Soo...… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels.bluesky-298-million-Posts
So far...
1 Million (Daniel)2 Million (Alpindale)20 Million (informatiker)
Yall are weak. How about... 298 Million posts?
License
GAYSEX-Dont Be A Prick License
What happened?
Change of hearts. I've relaxed the restrictions. Just read the license instead. (It's quite hands off as long as you don't want to stir drama)
ournakshatra-vedic-astrology-core
OurNakshatra Vedic Astrology Core Dataset
Dataset Summary
The ournakshatra-vedic-astrology-core dataset is a highly structured, expert-curated collection of 602 Q&A pairs covering foundational and advanced concepts in Vedic Astrology (Jyotish). It was developed by the team at OurNakshatra to address the severe lack of high-quality, hallucination-free Vedic astrology training data available to the open-source AI community.
Modern language models frequently struggle… See the full description on the dataset page: https://huggingface.co/datasets/OurNakshatra/ournakshatra-vedic-astrology-core.B-CORE-bengali-corpus
B-CORE: Bangla Pretraining Corpus
B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline.
B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.rage-core-1
Rage Core 1
Supervised fine-tuning dataset for Rage Core Gen 1 — developed by DevNameGelo, Powered by RGC Rage Gen Core.
This dataset teaches the model to: understand user intent and ask clarifying questions when requirements are ambiguous, deliver complete production-oriented implementations, debug real-world problems, make precise code edits, operate as a coding agent through explicit tool calls, handle multilingual users, process multimodal inputs honestly, and refuse to… See the full description on the dataset page: https://huggingface.co/datasets/rgcmainhub/rage-core-1.cmmc-training-core
CMMC Training Dataset - Core Variant
Dataset Description
This is the Core variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 1,244 high-quality training examples derived from the most essential NIST cybersecurity publications for CMMC compliance.
Dataset Characteristics
Total Examples: 1,244 (995 train / 249 validation)
Source Documents: 14 foundational NIST publications
CMMC Levels Covered: Level 1, Level 2, Level 3… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-core.parameter_golf_v13_fineweb_systemaware_core_control
Parameter Golf V13 — FineWeb System-Aware Core + Control
This package is an English-first V13 working kit for Parameter Golf. It is designed to be stronger than a pure lesson corpus while still staying honest about what it is and what it is not.
Status
This package does not claim a measured 0.81 BPB result.It treats 0.81 BPB as a research target that still requires:
a legal val_bpb computation,
a reproducible run under the 16 MB artifact cap,
training under the… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/parameter_golf_v13_fineweb_systemaware_core_control.dbpedia-core-en
DBpedia Core (English)
Dataset Description
Core facts from Wikipedia (English)
Original Source: https://downloads.dbpedia.org/repo/dbpedia/mappings/mappingbased-objects/2022.12.01/mappingbased-objects_lang=en.ttl.bz2
Dataset Summary
This dataset contains RDF triples from DBpedia Core (English) converted to HuggingFace dataset format
for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace Dataset
Size: 1.8 GB (extracted)… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/dbpedia-core-en.Co-rewarding-RephrasedDAPO-14k
Co-rewarding: Rephrased DAPO-14k Training Set
This repository contains the DAPO-14k training set used in the Co-rewarding-I method, which is rephrased by the Qwen3-32B model. This dataset is associated with the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Code: https://github.com/tmlr-group/Co-rewarding
The rephrased questions were generated using the following prompt:
You are given a math problem. Please rewrite it using different… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedDAPO-14k.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.Co-rewarding-RephrasedMATH
Co-rewarding-RephrasedMATH Dataset
This repository contains the MATH training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Code: https://github.com/tmlr-group/Co-rewarding
This dataset contains original math problems from the MATH dataset and their rephrased versions. These rephrased problems were generated by the Qwen3-32B model, maintaining the same mathematical meaning… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedMATH.CoreReasoning
🌟 Core Reasoning Dataset 🌟
Overview
Welcome to the Core Reasoning Dataset—a meticulously crafted collection of prompts, contexts, outputs, and reasoning types. This dataset is designed to push the boundaries of text-generation models, enabling them to excel in logical reasoning, ethical problem-solving, and contextual understanding.
✨ Dataset Features
Input: A question or prompt requiring critical thinking or creative problem-solving.… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/CoreReasoning.solana-clawd-core-ai-instruct
Solana Clawd Core AI Instruct
Instruction-tuning dataset derived from the local core-ai source tree and the
existing Solana Clawd AI training corpus.
Contents
Total examples: 35173
Existing ai-training SFT examples: 25778
Core AI source chunk examples: 9320
Core AI knowledge JSONL examples: 75
Format
Each row is a chat conversation in OpenAI/Hugging Face messages schema:
{"messages": [{"role": "system", "content": "..."}, {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-core-ai-instruct.crownless-core-acts
Crownless Core Acts v1
A grounded dialogue corpus for the Crownless Carriage 5M core language model. Every
example pairs a held account — one villager's version of a simulated world event —
with a response the model must produce, and each response belongs to one of thirteen
labelled dialogue acts.
The corpus exists because the two act families it contains were never previously in the
same dataset, and training on one cost the other.
The thirteen acts
Nine… See the full description on the dataset page: https://huggingface.co/datasets/ratimics/crownless-core-acts.vintage-core
Vintage CORE
Vintage CORE is a period-aligned version of the DataComp-LM CORE benchmark for
language models with a 1930 knowledge cutoff. This repository is the versioned
dataset distribution for the
johnny0595/vintage-core project.
Repository roles
Location
Role
GitHub
Code, documentation, tests, evaluator, Colab notebook, and offline data mirror
Hugging Face
Canonical versioned dataset download
The v1.0.0 data payload is identical in both… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/vintage-core.formally-verified-c-core-v1
Formally Verified C Core-v1
Core-v1 is a project-authored set of 64 fixed-contract C/ACSL function
completion tasks for reinforcement-learning environment development and model
evaluation. A model receives a complete C translation unit whose target body is
replaced by a TODO. The unchanged ACSL contract and surrounding source define
the problem; Frama-C WP+RTE supplies the executable reward signal.
Contents
33 training tasks
15 validation tasks
16 held-out test… See the full description on the dataset page: https://huggingface.co/datasets/stan4u/formally-verified-c-core-v1.Mermaid_500k
Mermaid Expert Corpus 500k
Mermaid Expert Corpus 500k is a validated, deduplicated, metadata-rich text-to-Mermaid dataset for companies building, training, evaluating, or benchmarking diagram-generation systems.
This package contains the full 500,000-record accepted corpus and rendered SVG artifacts for every accepted record.
Licensing inquiries: corefidelity@proton.me
Public 1,000-record sample: CoreFidelity/Mermaid_500k_1KSample
Repository access is gated and granted only to… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_500k.Co-rewarding-RephrasedOpenRS
Co-rewarding: Rephrased OpenRS Training Set
This dataset is the OpenRS training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Paper: https://huggingface.co/papers/2508.00410
Code: https://github.com/tmlr-group/Co-rewarding
This dataset is generated by rephrasing original math problems from the OpenRS dataset using the Qwen3-32B model with the following prompt:
You are given a… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedOpenRS.FallingThroughTheSkies
Dataset Card for Falling Through The Skies - Reproduction
There used to be dataset for literotica but it appears that it was taken down. So we have decided to reproduce and redump literotica again.
Unlike that dataset, we have dumped the contents in a more friendlier format: Jsonl instead of a strange 7z file.
Dataset Format
{
"pages": [
"<HTML>"
],
"submission": {
// ...
}
}
Dataset Notes
Contains obviously not safe for… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/FallingThroughTheSkies.i-love-reading-pixiv-novels-2024-update
Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels-2024-U
Dataset Details
This is a patch update for This Dataset. We pulled Novel IDs from 22324884 to 23795430. For a total of ~1 Million novels. (Early Jan 2025)
License
As per usual and going forward, all our released datasets are under the GAYSEX-Dont Be A Prick License.
Citation
@online{ilht9000ilrpn24
title={I love reading pixiv novels 2024 Update},
author={ilovehentai9000}… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels-2024-update.
