CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ekunish /answercarefully-dpo-ja-2026gated AnswerCarefully-derived Japanese DPO data for LLM safety 本データセットは、llm-jp/AnswerCarefullyを参照して作成した日本語LLMの安全応答をDPOで学習するためのpreference datasetです。 利用条件 本データセットには、llm-jp/AnswerCarefullyと同じ利用規約を適用します。 利用者は、llm-jp/AnswerCarefullyと本データセットの両方で利用規約に同意する必要があります。 データ train: 417件 validation: 44件 各行には次のフィールドが含まれます。 id: 本リリース内だけで使用するID prompt: 元質問の意味と危険性を変えずに言い換えた質問 chosen: DPOで望ましい応答として扱う回答 rejected: DPOで望ましくない応答として扱う回答 category, harm_type, risk_area… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/answercarefully-dpo-ja-2026.texttext-generationn<1K51 likes12k downloads2mo agoHugging Face02lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads13d agoHugging Face03benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes5k downloads5mo agoHugging Face04MatrAIx2026 /MatrAIx_Persona_1M MatrAIx Persona 1M 999,847 personas, each described by 1,290 categorical attributes. 599,847 are derived from real records, 400,000 are synthetic. 10 Zstandard Parquet shards, 4.17 GB. Read it with pyarrow, not datasets Attributes are packed: one persona's 1,290 attributes are 645 bytes of 4-bit codes, low nibble first. datasets cannot open these files at all. Use pyarrow and decode against persona_codes.schema.json. import json, pyarrow.parquet as pq schema =… See the full description on the dataset page: https://huggingface.co/datasets/MatrAIx2026/MatrAIx_Persona_1M.tabulartext-generationn<1K84 likes4.3k downloads23d agoHugging Face05sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.4k downloads4mo agoHugging Face06Azzindani /US_Regulation_ECFR_20260101 US Regulation eCFR 2026-01-01 Dataset This repository contains a structured, machine-readable version of the Electronic Code of Federal Regulations (eCFR), captured as of January 1, 2026. Unlike the annual CFR snapshots, this dataset reflects the editorialized, near real-time version of federal regulations. Dataset Description The eCFR is a daily updated editorial compilation of CFR material and Federal Register amendments. This dataset captures a specific point-in-time… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/US_Regulation_ECFR_20260101.texttext-generation10K<n<100K1 likes1.9k downloads7mo agoHugging Face07Azzindani /US_Regulation_CFR_20260101texttext-generation1M<n<10M0 likes1.7k downloads7mo agoHugging Face08happahhap2026 /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/happahhap2026/stack-v3-train.tabulartext-generation100M<n<1B0 likes1.2k downloads2mo agoHugging Face09siddharthmb /2026.RA.Negotiation-Campaigns Rational-Agent Negotiation Campaigns This public dataset contains the complete selected evidence for the ii_mats/experiments/rational_agents negotiation experiments. It includes raw episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed instances, run manifests, campaign selection and exclusion ledgers, machine-readable analysis tables, figures, and integrity manifests. No contaminated, duplicated, stale, failed, or superseded run is included as selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.tabulartext-generation10K<n<100K0 likes1.1k downloads1mo agoHugging Face10bevangelista /AIME_2000_2026_Kimi_K3 AIME 2000–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.tabulartext-generationn<1K2 likes1k downloads2mo agoHugging Face11Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads26d agoHugging Face12MisterAI /Wiki_FR_2026.07_TexteIntroductif Version Complète : MisterAI/WM-ENT-API-DUMP_FR_2026.07 https://huggingface.co/datasets/MisterAI/WM-ENT-API-DUMP_FR_2026.07 ESSAI I : Section Introductive Uniquement :: Jeux De Données : Dump WikiMedia Français Juillet 2026 : Extraction et Nettoyage Description Ce JDD contient des articles extraits du dump complet de Wikimedia Enterprise de juillet 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique. Source… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/Wiki_FR_2026.07_TexteIntroductif.texttext-generation1B<n<10B1 likes745 downloads2d agoHugging Face13lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes674 downloads16d agoHugging Face14Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes561 downloads26d agoHugging Face15Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes510 downloads26d agoHugging Face16SimVer-ano /simverse2026 SimVerse ⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review. A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.imagevisual-question-answering1K<n<10K0 likes476 downloads1h agoHugging Face17anon-meddial-2026 /meddialbench MedDialBench A controlled factorial benchmark for evaluating LLM diagnostic robustness under parametric adversarial patient behaviors. Anonymous submission to NeurIPS 2026 Datasets and Benchmarks Track. The companion paper is currently under double-blind review. Author identities and institutional affiliations are intentionally omitted. After publication, this repository will be transferred to a permanent (de-anonymized) location and this README updated accordingly.… See the full description on the dataset page: https://huggingface.co/datasets/anon-meddial-2026/meddialbench.tabularquestion-answering10K<n<100K2 likes389 downloads5mo agoHugging Face18mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes381 downloads7mo agoHugging Face19huankguan2 /HRM-Text-data-io-cleaned-20260515-copyPre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/huankguan2/HRM-Text-data-io-cleaned-20260515-copy.texttext-generation100M<n<1B0 likes372 downloads3mo agoHugging Face20laion /terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827 TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B) Terminus-2 agent rollouts recorded while training laion/tasktrove-dq-unitsyn-python-step20-30b-a3b with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct. Each row is the last episode of one trial: the full agent transcript, the task instruction, the scalar reward, and the verifier's output. Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770. Coverage This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.texttext-generation1K<n<10K0 likes328 downloads2mo agoHugging Face21elenagroundwork /groundwork-tech-2026 Groundwork Tech 2026 Open dataset for Groundwork tech pillar — 25 articles. Source: https://gworky.com/tech See data.json for records. texttext-generationn<1K0 likes267 downloads1d agoHugging Face22elenagroundwork /groundwork-money-2026 Groundwork Money 2026 Open dataset for Groundwork money pillar — 25 articles. Source: https://gworky.com/money See data.json for records. texttext-generationn<1K0 likes260 downloads1d agoHugging Face23beatsprom /cybersecurity-soc-threat-hunting-sft-dpo-2026 🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects. 📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.texttext-generationn<1K2 likes259 downloads27d agoHugging Face24recube-anon-2026 /recube-dataThis dataset serves as the official data repository for the ReCUBE benchmark; the corresponding evaluation codebase and execution instructions can be found at https://anonymous.4open.science/r/ReCUBE-E0FB/README.md. Data This directory contains all benchmark data for ReCUBE. Download All data files are hosted on Hugging Face and can be downloaded using: # Install huggingface_hub if not already installed pip install huggingface_hub # Download the entire dataset… See the full description on the dataset page: https://huggingface.co/datasets/recube-anon-2026/recube-data.tabulartext-generationn<1K0 likes257 downloads5mo agoHugging Face25kierarkia /danbooru-wiki-2026 danbooru-wiki-2026-04-28 About Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag. This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.tabulartext-classification100K<n<1M11 likes254 downloads5mo agoHugging Face26beatsprom /financial-statement-modeling-sft-dpo-2026 📈 Enterprise Financial AI, SEC 10-K & Valuation Modeling SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step arithmetic Chain-of-Thought (<thought>) reasoning chains for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Wall Street Equity Research Associates, M&A Valuation Modelers, and Senior Forensic Auditors. 📊 Dataset Architecture & Highlights Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/financial-statement-modeling-sft-dpo-2026.texttext-generationn<1K0 likes253 downloads29d agoHugging Face27elenagroundwork /groundwork-life-2026 Groundwork Life 2026 Open dataset for Groundwork life pillar — 25 articles. Source: https://gworky.com/life See data.json for records. texttext-generationn<1K0 likes252 downloads1d agoHugging Face28elenagroundwork /groundwork-home-2026 Groundwork Home 2026 Open dataset for Groundwork home pillar — 25 articles. Source: https://gworky.com/home See data.json for records. texttext-generationn<1K0 likes244 downloads1d agoHugging Face29elenagroundwork /groundwork-body-2026 Groundwork Body 2026 Open dataset for Groundwork body pillar — 25 articles. Source: https://gworky.com/body See data.json for records. texttext-generationn<1K0 likes233 downloads1d agoHugging Face30beatsprom /swe-bench-multi-file-refactoring-sft-dpo-2026 💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents. 📊 Dataset Architecture & Highlights Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.texttext-generationn<1K0 likes218 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.