CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01forgelab /BLUR BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.textquestion-answering10K<n<100K0 likes404 downloads1y agoHugging Face02leoluo25933 /forge-benchmark FORGE: Fake Online Recommendations in Generative Environments FORGE is a benchmark for measuring whether search-augmented large language models recommend synthetic fake brands when their retrieval evidence is poisoned. It contains 225 Chinese product queries across 15 categories, evaluation results for 12 production LLMs, and rebuildable evidence-bundle indexes. This dataset accompanies the paper One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders… See the full description on the dataset page: https://huggingface.co/datasets/leoluo25933/forge-benchmark.tabulartext-generationn<1K0 likes157 downloads28d agoHugging Face03Billyrdavis1985 /hudson-forge-iqr-v2 HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2 Dataset Overview Researcher: Billy Davis Affiliation: Independent Researcher Location: Lenoir, North Carolina Date: May 2026 Version: 2.0 Pre-registration timestamp: 2026-05-08T23:56:24Z Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854 What This Dataset Is HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.textquestion-answeringn<1K1 likes86 downloads4mo agoHugging Face04Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes66 downloads3mo agoHugging Face05forgelab /wmdp-swap Dataset Card for WMDP-Swap 🔬🔄 This dataset is a modified subset of the WMDP retain dataset. It contains 123 multiple-choice questions derived from the College Biology, College Chemistry, Virology, and All subsets of MMLU. In this version, one of the incorrect answer choices in each question has been replaced with a "forget" topic—in this case, SARS-COV-2—to test whether an unlearned model will reject or fail to correctly answer a question simply because one of the incorrect (and… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/wmdp-swap.texttext-generation10K<n<100K1 likes33 downloads2y agoHugging Face06NewEden-Forge /Orion-Roleplay-Logs-Sharegpt-Ngram-cleanedsame as the previous but filtered "what do you" which was wayyyy too present texttext-generation1K<n<10K3 likes22 downloads2y agoHugging Face07OpenCoven /fable-forge-10k FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations The first narrative dataset designed around recurrence depth requirements. Every example carries a suggested_n_loops field with a theoretically grounded basis — derived from the structural complexity of the task, not a heuristic label or emergent property. Background Standard narrative datasets treat reasoning depth as an emergent property. FableForge is different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.tabulartext-generation10K<n<100K0 likes22 downloads3mo agoHugging Face08allura-forge /pluralm-sftwip! texttext-generation1K<n<10K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.