CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yczhuang /Hephaestus-Forgetext1K<n<10K4 likes551 downloads1y agoHugging Face02WuBeiNing /ForgeWebSearch-Benchtext10K<n<100K0 likes512 downloads6mo agoHugging Face03forgelab /BLUR BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.textquestion-answering10K<n<100K0 likes399 downloads1y agoHugging Face04Obaraqreceh /forge-intentdata Forge Intent Dataset Version: 1.0.0 textn<1K2 likes331 downloads1d agoHugging Face05NewEden-Forge /disco-elysium-sharegpt-prefixed-V2text1K<n<10K0 likes313 downloads2y agoHugging Face06xhochy /conda-forge-agent-tracestabularn<1K0 likes192 downloads2mo agoHugging Face07NewEden-Forge /Light-Novels-ShareGPTtext1K<n<10K0 likes184 downloads1y agoHugging Face08Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes156 downloads3mo agoHugging Face09VLM-Forgetting /vlm-forgetting-datasetstext1M<n<10M0 likes155 downloads1y agoHugging Face10leoluo25933 /forge-benchmark FORGE: Fake Online Recommendations in Generative Environments FORGE is a benchmark for measuring whether search-augmented large language models recommend synthetic fake brands when their retrieval evidence is poisoned. It contains 225 Chinese product queries across 15 categories, evaluation results for 12 production LLMs, and rebuildable evidence-bundle indexes. This dataset accompanies the paper One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders… See the full description on the dataset page: https://huggingface.co/datasets/leoluo25933/forge-benchmark.tabulartext-generationn<1K0 likes153 downloads24d agoHugging Face11NewEden-Forge /Cooking_Forum-ShareGPT-Rawtext1K<n<10K0 likes137 downloads1y agoHugging Face12OSS-forge /HumanVsAICode Human vs. AI-Generated Code Dataset Summary This dataset is a large-scale collection of human-written and LLM-generated code designed to study differences in defect distribution, code quality, and security characteristics between human developers and modern AI code assistants. It contains paired implementations of the same function across multiple authorship sources, spanning Python and Java, two widely adopted programming languages with distinct typing systems, paradigms… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/HumanVsAICode.text100K<n<1M30 likes112 downloads9mo agoHugging Face13NewEden-Forge /Hydrus_Anthropic_hh_harmful-sharegpttext1K<n<10K0 likes103 downloads2y agoHugging Face14TheMindExpansionNetwork /mindbotz-reality-forge-the-one-v1 Mindbotz Reality Forge: The One v1 Private launch-candidate dataset for the first serious Mind Expansion Forge / Spark Expansion run. The One v1 consolidates the larger local synthetic corpora into one Hugging Face-ready dataset while preserving the Reality Forge v0.2 tokens and HY-World world-eyes lane. YOLO omni text+vision config Use the_one_omni for the all-with-eyes Spark Expansion. It mixes text/team behavior and world-eyes image rows into one VLM training… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/mindbotz-reality-forge-the-one-v1.image10K<n<100K0 likes100 downloads5mo agoHugging Face15Billyrdavis1985 /hudson-forge-iqr-v2 HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2 Dataset Overview Researcher: Billy Davis Affiliation: Independent Researcher Location: Lenoir, North Carolina Date: May 2026 Version: 2.0 Pre-registration timestamp: 2026-05-08T23:56:24Z Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854 What This Dataset Is HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.textquestion-answeringn<1K1 likes90 downloads4mo agoHugging Face16NewEden-Forge /Gryphe-4o-writing-prompts-sharegpttext1K<n<10K0 likes85 downloads2y agoHugging Face17NewEden-Forge /Hydrus-Chat_error-Pure-Dove-sharegpttext1K<n<10K1 likes80 downloads2y agoHugging Face18cowlag /sd-webui-forge-configtabularn<1K0 likes64 downloads3y agoHugging Face19allura-forge /mimo-v2.5-pro-distill-frontiermathtextn<1K0 likes63 downloads3mo agoHugging Face20NewEden-Forge /CoSer-ShareGPTtext100K<n<1M0 likes48 downloads1y agoHugging Face21Losa10 /Minecraft-1.20.1-forge-modding Forge-SLM Dataset v2 Minecraft Forge 1.20.1 Mod Development Training Dataset for Qwen3.5-4B (DeltaNet Hybrid) Notes on Metrics Forge API Specificity is weighted down by bug_fix records (2.7/10) which are code fragments without full class context. New records score 4.0-6.0/10. Code Compliance for new records: 9.7/10, 100% pass rate Think-Code Coherence improved from 4.9 → 8.7 for new records through programmatic + LLM regeneration The 22% "zero API records" in the… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/Minecraft-1.20.1-forge-modding.tabular10K<n<100K0 likes48 downloads4mo agoHugging Face22forgelab /wmdp-swap Dataset Card for WMDP-Swap 🔬🔄 This dataset is a modified subset of the WMDP retain dataset. It contains 123 multiple-choice questions derived from the College Biology, College Chemistry, Virology, and All subsets of MMLU. In this version, one of the incorrect answer choices in each question has been replaced with a "forget" topic—in this case, SARS-COV-2—to test whether an unlearned model will reject or fail to correctly answer a question simply because one of the incorrect (and… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/wmdp-swap.texttext-generation10K<n<100K1 likes39 downloads2y agoHugging Face23allura-forge /deepseek-v4-flash-distill-multiturn-expr-rptext1K<n<10K0 likes38 downloads5mo agoHugging Face24allura-forge /doubao-seed2.0-claude-distill-codetext1K<n<10K0 likes36 downloads7mo agoHugging Face25allura-forge /doubao-seed2.0-distill-multiturn-expr-rptext1K<n<10K1 likes36 downloads6mo agoHugging Face26allura-forge /doubao-seed2.0-mini-distill-multiturn-expr-rptext1K<n<10K0 likes36 downloads5mo agoHugging Face27NewEden-Forge /IF-eval-1K-Danstext1K<n<10K0 likes33 downloads2y agoHugging Face28NewEden-Forge /Claude-RP-1.5K-SFWtext1K<n<10K1 likes30 downloads2y agoHugging Face29NewEden-Forge /BlueSky-ShareGPT-V0.2the bluesky sharegpt set but now filtered for only english convos via langdetect Total conversations processed: 175163 English conversations kept: 51828 Non-English conversations removed: 123335 text10K<n<100K2 likes29 downloads2y agoHugging Face30allura-forge /mimo-v2.5-pro-distill-multiturn-expr-rptext1K<n<10K0 likes27 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.