datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hephaestus-ForgeForgeWebSearch-BenchBLUR
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.forge-intentdata
Forge Intent Dataset
Version: 1.0.0
disco-elysium-sharegpt-prefixed-V2conda-forge-agent-tracesLight-Novels-ShareGPTforge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.vlm-forgetting-datasetsforge-benchmark
FORGE: Fake Online Recommendations in Generative Environments
FORGE is a benchmark for measuring whether search-augmented large language
models recommend synthetic fake brands when their retrieval evidence is
poisoned. It contains 225 Chinese product queries across 15 categories,
evaluation results for 12 production LLMs, and rebuildable evidence-bundle
indexes.
This dataset accompanies the paper One Polluted Page Is Enough: Evaluating
Web Content Pollution in LLM Recommenders… See the full description on the dataset page: https://huggingface.co/datasets/leoluo25933/forge-benchmark.Cooking_Forum-ShareGPT-RawHumanVsAICode
Human vs. AI-Generated Code
Dataset Summary
This dataset is a large-scale collection of human-written and LLM-generated code designed to study differences in defect distribution, code quality, and security characteristics between human developers and modern AI code assistants.
It contains paired implementations of the same function across multiple authorship sources, spanning Python and Java, two widely adopted programming languages with distinct typing systems, paradigms… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/HumanVsAICode.Hydrus_Anthropic_hh_harmful-sharegptmindbotz-reality-forge-the-one-v1
Mindbotz Reality Forge: The One v1
Private launch-candidate dataset for the first serious Mind Expansion Forge / Spark Expansion run.
The One v1 consolidates the larger local synthetic corpora into one Hugging Face-ready dataset while preserving the Reality Forge v0.2 tokens and HY-World world-eyes lane.
YOLO omni text+vision config
Use the_one_omni for the all-with-eyes Spark Expansion. It mixes text/team behavior and world-eyes image rows into one VLM training… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/mindbotz-reality-forge-the-one-v1.hudson-forge-iqr-v2
HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2
Dataset Overview
Researcher: Billy Davis
Affiliation: Independent Researcher
Location: Lenoir, North Carolina
Date: May 2026
Version: 2.0
Pre-registration timestamp: 2026-05-08T23:56:24Z
Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854
What This Dataset Is
HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.Gryphe-4o-writing-prompts-sharegptHydrus-Chat_error-Pure-Dove-sharegptsd-webui-forge-configmimo-v2.5-pro-distill-frontiermathCoSer-ShareGPTMinecraft-1.20.1-forge-modding
Forge-SLM Dataset v2
Minecraft Forge 1.20.1 Mod Development Training Dataset for Qwen3.5-4B (DeltaNet Hybrid)
Notes on Metrics
Forge API Specificity is weighted down by bug_fix records (2.7/10) which are code fragments without full class context. New records score 4.0-6.0/10.
Code Compliance for new records: 9.7/10, 100% pass rate
Think-Code Coherence improved from 4.9 → 8.7 for new records through programmatic + LLM regeneration
The 22% "zero API records" in the… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/Minecraft-1.20.1-forge-modding.wmdp-swap
Dataset Card for WMDP-Swap 🔬🔄
This dataset is a modified subset of the WMDP retain dataset. It contains 123 multiple-choice questions derived from the College Biology, College Chemistry, Virology, and All subsets of MMLU. In this version, one of the incorrect answer choices in each question has been replaced with a "forget" topic—in this case, SARS-COV-2—to test whether an unlearned model will reject or fail to correctly answer a question simply because one of the incorrect (and… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/wmdp-swap.deepseek-v4-flash-distill-multiturn-expr-rpdoubao-seed2.0-claude-distill-codedoubao-seed2.0-distill-multiturn-expr-rpdoubao-seed2.0-mini-distill-multiturn-expr-rpIF-eval-1K-DansClaude-RP-1.5K-SFWBlueSky-ShareGPT-V0.2the bluesky sharegpt set but now filtered for only english convos via langdetect
Total conversations processed: 175163
English conversations kept: 51828
Non-English conversations removed: 123335
mimo-v2.5-pro-distill-multiturn-expr-rp
