CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reasoning-core /procedural-pile Procedural Pile: Procedural reasoning data SFT (and RL) Procedural Pile is a synthetic corpus of verifiable reasoning problems generated by Reasoning Core. It is intended for continued pretraining, mid-training, and supervised fine-tuning. Answers come from procedural generators and task-specific solvers or checkers, rather than language-model generation. The corpus spans mathematics, formal logic, planning, graphs, parsing, code, structured data, and other symbolic domains.… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-core/procedural-pile.texttext-generation10M<n<100M9 likes1.6k downloads7h agoHugging Face02agentlans /DSULT-Core-ShareGPT-X DSULT-Core/ShareGPT-X Filtered Dataset This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines. The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.tabulartext-generation10K<n<100K3 likes784 downloads10mo agoHugging Face03Lots-of-LoRAs /task1390_wscfixed_coreference Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1390_wscfixed_coreference Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1390_wscfixed_coreference.texttext-generationn<1K0 likes716 downloads2y agoHugging Face04Lots-of-LoRAs /task891_gap_coreference_resolution Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task891_gap_coreference_resolution Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task891_gap_coreference_resolution.texttext-generationn<1K0 likes559 downloads2y agoHugging Face05reasoning-core /symbolic-reasoning-env ⚠️DEPRECATED: PLEASE MOVE to hf.co/reasoning-core/procedural-pile Reasoning Core ◉ Paper: Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning Code: GitHub Repository reasoning-core is a text-based RLVR for LLM reasoning training. It is centered on expressive symbolic tasks, including full fledged FOL, formal mathematics with TPTP, formal planning with novel domains, and syntax tasks. Abstract We introduce Reasoning Core, a new… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-core/symbolic-reasoning-env.texttext-generation100K<n<1M14 likes504 downloads7h agoHugging Face06Lots-of-LoRAs /task893_gap_fill_the_blank_coreference_resolution Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task893_gap_fill_the_blank_coreference_resolution Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task893_gap_fill_the_blank_coreference_resolution.texttext-generationn<1K0 likes409 downloads2y agoHugging Face07Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes298 downloads1mo agoHugging Face08DSULT-Core /ShareGPT-X Dataset Summary ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed. Supported Tasks and Leaderboards text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.tabulartext-generation100K<n<1M17 likes222 downloads1y agoHugging Face09laion /COREX-18textCORE-18 Fulltext Introducing the CORE-18 Full Text dataset, among the first well-maintained public datasets of CORE. CORE offers one of the largest collections of research papers, including supplementary metadata, to support Artificial Intelligence, Machine Learning research, and engineering projects. This dataset has gained significant attention among major corporations and research laboratories for Natural Language Processing research. Recognizing the importance of accessibility… See the full description on the dataset page: https://huggingface.co/datasets/laion/COREX-18text.translation1 likes210 downloads2y agoHugging Face10Lots-of-LoRAs /task892_gap_reverse_coreference_resolution Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task892_gap_reverse_coreference_resolution Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task892_gap_reverse_coreference_resolution.texttext-generationn<1K0 likes161 downloads2y agoHugging Face11DSULT-Core /i-love-reading-pixiv-novels Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels Dataset Details Dataset Description This is a more or less raw dump of pixiv novel data (13,012,017 documents to be exact.) Are you the hacker? I scraped pixiv on the same day of the Kadokawa site issues. I had no clue about the issue surrounding nicolive, etc until I noticed after the scrape was done. around 8 hours before I started the scrape, the websites(?) went down. Soo...… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels.text-generation10M<n<100M8 likes121 downloads2y agoHugging Face12DSULT-Core /bluesky-298-million-Posts So far... 1 Million (Daniel)2 Million (Alpindale)20 Million (informatiker) Yall are weak. How about... 298 Million posts? License GAYSEX-Dont Be A Prick License What happened? Change of hearts. I've relaxed the restrictions. Just read the license instead. (It's quite hands off as long as you don't want to stir drama) text-generation50 likes118 downloads2y agoHugging Face13OurNakshatra /ournakshatra-vedic-astrology-core OurNakshatra Vedic Astrology Core Dataset Dataset Summary The ournakshatra-vedic-astrology-core dataset is a highly structured, expert-curated collection of 602 Q&A pairs covering foundational and advanced concepts in Vedic Astrology (Jyotish). It was developed by the team at OurNakshatra to address the severe lack of high-quality, hallucination-free Vedic astrology training data available to the open-source AI community. Modern language models frequently struggle… See the full description on the dataset page: https://huggingface.co/datasets/OurNakshatra/ournakshatra-vedic-astrology-core.textquestion-answeringn<1K2 likes105 downloads3mo agoHugging Face14nahid-hub /B-CORE-bengali-corpus B-CORE: Bangla Pretraining Corpus B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline. B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.texttext-generation10M<n<100M0 likes97 downloads2mo agoHugging Face15rgcmainhub /rage-core-1 Rage Core 1 Supervised fine-tuning dataset for Rage Core Gen 1 — developed by DevNameGelo, Powered by RGC Rage Gen Core. This dataset teaches the model to: understand user intent and ask clarifying questions when requirements are ambiguous, deliver complete production-oriented implementations, debug real-world problems, make precise code edits, operate as a coding agent through explicit tool calls, handle multilingual users, process multimodal inputs honestly, and refuse to… See the full description on the dataset page: https://huggingface.co/datasets/rgcmainhub/rage-core-1.texttext-generationn<1K0 likes83 downloads14d agoHugging Face16ethanolivertroy /cmmc-training-core CMMC Training Dataset - Core Variant Dataset Description This is the Core variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 1,244 high-quality training examples derived from the most essential NIST cybersecurity publications for CMMC compliance. Dataset Characteristics Total Examples: 1,244 (995 train / 249 validation) Source Documents: 14 foundational NIST publications CMMC Levels Covered: Level 1, Level 2, Level 3… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-core.texttext-generation1K<n<10K1 likes71 downloads11mo agoHugging Face178Planetterraforming /parameter_golf_v13_fineweb_systemaware_core_control Parameter Golf V13 — FineWeb System-Aware Core + Control This package is an English-first V13 working kit for Parameter Golf. It is designed to be stronger than a pure lesson corpus while still staying honest about what it is and what it is not. Status This package does not claim a measured 0.81 BPB result.It treats 0.81 BPB as a research target that still requires: a legal val_bpb computation, a reproducible run under the 16 MB artifact cap, training under the… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/parameter_golf_v13_fineweb_systemaware_core_control.text-generationn<1K1 likes71 downloads5mo agoHugging Face18CleverThis /dbpedia-core-en DBpedia Core (English) Dataset Description Core facts from Wikipedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/mappings/mappingbased-objects/2022.12.01/mappingbased-objects_lang=en.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia Core (English) converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 1.8 GB (extracted)… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/dbpedia-core-en.texttext-generation10M<n<100M0 likes63 downloads10mo agoHugging Face19TMLR-Group-HF /Co-rewarding-RephrasedDAPO-14k Co-rewarding: Rephrased DAPO-14k Training Set This repository contains the DAPO-14k training set used in the Co-rewarding-I method, which is rephrased by the Qwen3-32B model. This dataset is associated with the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models. Code: https://github.com/tmlr-group/Co-rewarding The rephrased questions were generated using the following prompt: You are given a math problem. Please rewrite it using different… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedDAPO-14k.texttext-generation10K<n<100K0 likes61 downloads1y agoHugging Face20aakashMeghwar01 /Sindhi-Intelligence-Core-SFT 🧠 Sindhi Intelligence Core SFT This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning. 📊 Dataset Summary This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT). 📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.texttext-generation100K<n<1M1 likes57 downloads7mo agoHugging Face21TMLR-Group-HF /Co-rewarding-RephrasedMATH Co-rewarding-RephrasedMATH Dataset This repository contains the MATH training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models. Code: https://github.com/tmlr-group/Co-rewarding This dataset contains original math problems from the MATH dataset and their rephrased versions. These rephrased problems were generated by the Qwen3-32B model, maintaining the same mathematical meaning… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedMATH.texttext-generation1K<n<10K0 likes51 downloads1y agoHugging Face22ayjays132 /CoreReasoning 🌟 Core Reasoning Dataset 🌟 Overview Welcome to the Core Reasoning Dataset—a meticulously crafted collection of prompts, contexts, outputs, and reasoning types. This dataset is designed to push the boundaries of text-generation models, enabling them to excel in logical reasoning, ethical problem-solving, and contextual understanding. ✨ Dataset Features Input: A question or prompt requiring critical thinking or creative problem-solving.… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/CoreReasoning.texttext-generation1K<n<10K2 likes45 downloads2y agoHugging Face23solanaclawd /solana-clawd-core-ai-instruct Solana Clawd Core AI Instruct Instruction-tuning dataset derived from the local core-ai source tree and the existing Solana Clawd AI training corpus. Contents Total examples: 35173 Existing ai-training SFT examples: 25778 Core AI source chunk examples: 9320 Core AI knowledge JSONL examples: 75 Format Each row is a chat conversation in OpenAI/Hugging Face messages schema: {"messages": [{"role": "system", "content": "..."}, {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-core-ai-instruct.texttext-generation10K<n<100K0 likes45 downloads3mo agoHugging Face24ratimics /crownless-core-acts Crownless Core Acts v1 A grounded dialogue corpus for the Crownless Carriage 5M core language model. Every example pairs a held account — one villager's version of a simulated world event — with a response the model must produce, and each response belongs to one of thirteen labelled dialogue acts. The corpus exists because the two act families it contains were never previously in the same dataset, and training on one cost the other. The thirteen acts Nine… See the full description on the dataset page: https://huggingface.co/datasets/ratimics/crownless-core-acts.text-generation0 likes43 downloads10d agoHugging Face25jbduran /vintage-core Vintage CORE Vintage CORE is a period-aligned version of the DataComp-LM CORE benchmark for language models with a 1930 knowledge cutoff. This repository is the versioned dataset distribution for the johnny0595/vintage-core project. Repository roles Location Role GitHub Code, documentation, tests, evaluator, Colab notebook, and offline data mirror Hugging Face Canonical versioned dataset download The v1.0.0 data payload is identical in both… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/vintage-core.text-generation1 likes35 downloads1mo agoHugging Face26stan4u /formally-verified-c-core-v1 Formally Verified C Core-v1 Core-v1 is a project-authored set of 64 fixed-contract C/ACSL function completion tasks for reinforcement-learning environment development and model evaluation. A model receives a complete C translation unit whose target body is replaced by a TODO. The unchanged ACSL contract and surrounding source define the problem; Frama-C WP+RTE supplies the executable reward signal. Contents 33 training tasks 15 validation tasks 16 held-out test… See the full description on the dataset page: https://huggingface.co/datasets/stan4u/formally-verified-c-core-v1.text-generationn<1K0 likes34 downloads14d agoHugging Face27CoreFidelity /Mermaid_500kgated Mermaid Expert Corpus 500k Mermaid Expert Corpus 500k is a validated, deduplicated, metadata-rich text-to-Mermaid dataset for companies building, training, evaluating, or benchmarking diagram-generation systems. This package contains the full 500,000-record accepted corpus and rendered SVG artifacts for every accepted record. Licensing inquiries: corefidelity@proton.me Public 1,000-record sample: CoreFidelity/Mermaid_500k_1KSample Repository access is gated and granted only to… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_500k.text-generation100K<n<1M0 likes33 downloads4mo agoHugging Face28TMLR-Group-HF /Co-rewarding-RephrasedOpenRS Co-rewarding: Rephrased OpenRS Training Set This dataset is the OpenRS training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models. Paper: https://huggingface.co/papers/2508.00410 Code: https://github.com/tmlr-group/Co-rewarding This dataset is generated by rephrasing original math problems from the OpenRS dataset using the Qwen3-32B model with the following prompt: You are given a… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedOpenRS.texttext-generation1K<n<10K0 likes32 downloads1y agoHugging Face29DSULT-Core /FallingThroughTheSkies Dataset Card for Falling Through The Skies - Reproduction There used to be dataset for literotica but it appears that it was taken down. So we have decided to reproduce and redump literotica again. Unlike that dataset, we have dumped the contents in a more friendlier format: Jsonl instead of a strange 7z file. Dataset Format { "pages": [ "<HTML>" ], "submission": { // ... } } Dataset Notes Contains obviously not safe for… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/FallingThroughTheSkies.text-generation1 likes28 downloads2y agoHugging Face30DSULT-Core /i-love-reading-pixiv-novels-2024-update Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels-2024-U Dataset Details This is a patch update for This Dataset. We pulled Novel IDs from 22324884 to 23795430. For a total of ~1 Million novels. (Early Jan 2025) License As per usual and going forward, all our released datasets are under the GAYSEX-Dont Be A Prick License. Citation @online{ilht9000ilrpn24 title={I love reading pixiv novels 2024 Update}, author={ilovehentai9000}… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels-2024-update.text-generation1M<n<10M2 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.