CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /Ling-Coder-SFT 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.texttext-generation1M<n<10M45 likes1.4k downloads1y agoHugging Face02Emulated-Inc /forum-competition-math-training-pool Forum competition mathematics training pool Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.texttext-generation100K<n<1M0 likes741 downloads11d agoHugging Face03inclusionAI /AReaL-tau2-data AReaL-tau2-data Synthetic training data for multi-turn interactive tool-using agents, generated by SEA, a self-evolving multi-agent data engine. This dataset is used to train AReaL-SEA-235B-A22B, achieving state-of-the-art results on τ²-bench. Paper: From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents Training Framework: AReaL Benchmark: τ²-bench Dataset Overview The dataset covers three customer-service… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AReaL-tau2-data.text-generation10K<n<100K15 likes706 downloads7mo agoHugging Face04Emulated-Inc /olympiad-math-training-pool Olympiad mathematics training pool Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.texttext-generation100K<n<1M0 likes688 downloads11d agoHugging Face05inclusionAI /SWE-CARE SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation Dataset Description SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.texttext-generation1K<n<10K8 likes581 downloads11mo agoHugging Face06zhush /incantation-elden-ring-scenes Incantation Elden Ring Combat Captions Paper | Project page | GitHub Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper. This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.textvideo-classification1K<n<10K3 likes517 downloads2mo agoHugging Face07hheiden /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B8 likes501 downloads8mo agoHugging Face08inclusionAI /Ling-Coder-SyntheticQA 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.texttext-generation10M<n<100M17 likes479 downloads1y agoHugging Face09Emulated-Inc /competition-math-training-pool Competition mathematics training pool Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.texttext-generation1M<n<10M0 likes421 downloads12d agoHugging Face10Parakeet-Inc /joyo-kanji-yomi-benchmark-parakeet 日本語 | English 常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet) 常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。 このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。 このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。 概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.texttext-to-speech10K<n<100K5 likes394 downloads1mo agoHugging Face11Bilsteen /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes383 downloads1mo agoHugging Face12inclusionAI /A3S-Bench Agent3Sigma-Stage (A3S-Bench) 💻 GitHub | 🏆 Leaderboard | 📄 Paper (PDF) | arXiv Agent3Sigma-Stage (A3S-Bench) is an end-to-end security evaluation framework for autonomous agents (e.g., OpenClaw), designed to systematically measure both an Agent's ability to resist attacks during multi-turn interactions and its utility in completing legitimate tasks. The framework provides an evaluation dataset covering 10 security risk categories across 6 real-world usage scenarios… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/A3S-Bench.texttext-generation1K<n<10K2 likes291 downloads4mo agoHugging Face13simplex-ai-inc /LiteResearcher-SFT-Data LiteResearcher — SFT Cold-Start Data Distilled deep-research trajectories used for the SFT cold-start of LiteResearcher-4B This dataset contains the 68,231 multi-turn deep-research trajectories used to train the SFT cold-start checkpoint that RL (GRPO+TIS) is later launched from — the "68.2 K distilled deep-research trajectories" referenced in the paper and in LiteResearcher-Data. Each row is a complete ReAct-style episode: a research question, the model's interleaved… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-SFT-Data.textquestion-answering10K<n<100K0 likes270 downloads2mo agoHugging Face14Lots-of-LoRAs /task070_abductivenli_incorrect_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task070_abductivenli_incorrect_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task070_abductivenli_incorrect_classification.texttext-generation1K<n<10K0 likes262 downloads2y agoHugging Face15inclusionAI /Ring-lite-sft-data 🤖 ModelScope 🤗 HuggingFace 🖥️ GitHub Ring-lite-sft-data This is a the SFT data used during the fine-tuning of the Ring-lite model. The query pool was sourced from open-source repositories and further enriched through synthetic generation using large language models (LLMs). To ensure the production of high-fidelity responses with Long-CoT, we implemented an iterative refinement pipeline that synergistically combines automated model… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-sft-data.texttext-generation1M<n<10M14 likes248 downloads1y agoHugging Face16Lots-of-LoRAs /task454_swag_incorrect_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task454_swag_incorrect_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task454_swag_incorrect_answer_generation.texttext-generation1K<n<10K0 likes224 downloads2y agoHugging Face17Emulated-Inc /logic-grid-puzzles-training-pool Logic grid puzzles training pool Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a list of clues that together admit exactly one arrangement. Two sets drawn for this pool by generators run here under the seeds recorded below, and two public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.texttext-generation100K<n<1M0 likes169 downloads11d agoHugging Face18Lots-of-LoRAs /task068_abductivenli_incorrect_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task068_abductivenli_incorrect_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task068_abductivenli_incorrect_answer_generation.texttext-generation1K<n<10K0 likes157 downloads2y agoHugging Face19Emulated-Inc /python-unit-test-training-pool Python unit test training pool A pool of public data for training a model to write tests for Python code. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named, and the only rows removed are the ones an overlap filter flagged against held-out material this pool is kept separate from. Every row of the normalised layer pairs a program with tests for it. That is the point of the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.texttext-generation1M<n<10M0 likes148 downloads11d agoHugging Face20th-laurel /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes147 downloads6mo agoHugging Face21Emulated-Inc /code-execution-trace-training-pool Code execution trace training pool Public Python code paired with one concrete call and the value that call returns. Every value in this pool was computed by running the code, not copied from a label. The data is laid out twice, and either layer may be used. pool/ Every source rewritten into one shape, 2174322 rows over 11 gzipped parts, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this pool code… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/code-execution-trace-training-pool.texttext-generation1M<n<10M1 likes144 downloads11d agoHugging Face22Emulated-Inc /symbolic-calculus-training-pool Symbolic calculus training pool Single variable calculus exercises with closed form answers: derivatives of composite expressions, indefinite and definite integrals, limits of indeterminate forms, and coefficients of Maclaurin series, with some multivariable operators in one of the sources. A set generated for this pool and two public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/symbolic-calculus-training-pool.texttext-generation1K<n<10K2 likes135 downloads11d agoHugging Face23Emulated-Inc /constrained-instruction-training-pool Constrained instruction training pool Public prompts for writing tasks, many of them carrying a constraint a program can check, from ten datasets read at the pinned revisions named below and one layer built here from them. The pool is laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 462652 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/constrained-instruction-training-pool.texttext-generation100K<n<1M0 likes132 downloads11d agoHugging Face24MatrixTeam /incantation-elden-ring-scenes Incantation Elden Ring Combat Captions Preview subset. This repository is an early public preview and reference subset of the Incantation dataset. It is provided to document the data format, annotation style, and initial training material ahead of the full paper/project release. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper. This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/MatrixTeam/incantation-elden-ring-scenes.textvideo-classification1K<n<10K1 likes129 downloads4mo agoHugging Face25Lots-of-LoRAs /task043_essential_terms_answering_incomplete_questions Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task043_essential_terms_answering_incomplete_questions Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task043_essential_terms_answering_incomplete_questions.texttext-generation1K<n<10K0 likes128 downloads2y agoHugging Face26Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes125 downloads10d agoHugging Face27WhySoCodius /in-context-grid-reasoning In-Context Grid Reasoning (ICGR) A small, fully synthetic benchmark for demonstration-conditioned rule induction: each task shows 2–4 (input grid → output grid) support pairs that share one hidden transformation, and the model must apply the same transformation to a held-out query input. It targets the same behaviour probed by recent in-context / latent-reasoning work on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning, arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.tabulartext-generation1K<n<10K1 likes122 downloads20d agoHugging Face28Lots-of-LoRAs /task297_storycloze_incorrect_end_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task297_storycloze_incorrect_end_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task297_storycloze_incorrect_end_classification.texttext-generation1K<n<10K0 likes115 downloads2y agoHugging Face29Emulated-Inc /python-functions-training-pool Python function-writing training pool A pool of public data for training a model to write Python functions. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named, and the only rows removed are the ones an overlap filter flagged against held-out material this pool is kept separate from. Rows in the normalised layer: 5756045. Rows in the raw layer: 6258415. The two layers pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.texttext-generation1M<n<10M0 likes111 downloads12d agoHugging Face30inclusionAI /Ring-lite-rl-data 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ring-lite-rl-data This dataset is a curated subset of high-quality problems across mathematics and code domains designed for reinforcement learning in the Ring-lite model. This dataset contains: Mathematics: Over 39,000 rigorously curated problems sourced from: Open-source datasets (BigMath, DeepScaleR, DAPO, DeepMath-103K) Art of Problem Solving (AoPS) contest collections Code: Approximately 8,400… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-rl-data.texttext-generation10K<n<100K8 likes110 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.