datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Omega-Forge-1M
Open-Omega-Forge-1M
Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.BLUR
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.forge-benchmark
FORGE: Fake Online Recommendations in Generative Environments
FORGE is a benchmark for measuring whether search-augmented large language
models recommend synthetic fake brands when their retrieval evidence is
poisoned. It contains 225 Chinese product queries across 15 categories,
evaluation results for 12 production LLMs, and rebuildable evidence-bundle
indexes.
This dataset accompanies the paper One Polluted Page Is Enough: Evaluating
Web Content Pollution in LLM Recommenders… See the full description on the dataset page: https://huggingface.co/datasets/leoluo25933/forge-benchmark.sigil-forge-training
SIGIL Forge Training Data
Forge-verified training tasks, references, fixtures, and versioned MLX SFT corpora for SIGIL.
The SIGIL source repository pins immutable revisions and verifies MANIFEST.json plus every payload.
Evaluation tasks and validation records are intentionally stored in a separate private repository.
mbpp-code-rl
MBPP for code RL (deduplicated against MBPP+)
MBPP prepared for RLVR training in verl,
with two independent hold-outs so both MBPP+ and MBPP's own canonical test
split stay reportable after training on this data.
split
rows
contents
train
320
MBPP canonical train + validation + prompt, minus everything in MBPP+
test
378
exactly the problems in evalplus/mbppplus
heldout_mbpp_test
276
MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.hudson-forge-iqr-v2
HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2
Dataset Overview
Researcher: Billy Davis
Affiliation: Independent Researcher
Location: Lenoir, North Carolina
Date: May 2026
Version: 2.0
Pre-registration timestamp: 2026-05-08T23:56:24Z
Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854
What This Dataset Is
HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.wmdp-swap
Dataset Card for WMDP-Swap 🔬🔄
This dataset is a modified subset of the WMDP retain dataset. It contains 123 multiple-choice questions derived from the College Biology, College Chemistry, Virology, and All subsets of MMLU. In this version, one of the incorrect answer choices in each question has been replaced with a "forget" topic—in this case, SARS-COV-2—to test whether an unlearned model will reject or fail to correctly answer a question simply because one of the incorrect (and… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/wmdp-swap.ParallelPrompt
PARALLELPROMPT
A benchmark dataset of 37,021 parallelizable prompts from real-world LLM conversations, designed for optimizing LLM serving systems through intra-query parallelism.
Repository and Resources
Dataset: Hugging Face
Code: GitHub
Paper: PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
The GitHub repository contains:
Data curation pipeline
Schema extraction code
Evaluation suite for measuring latency and quality
Baseline implementations… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/ParallelPrompt.Orion-Roleplay-Logs-Sharegpt-Ngram-cleanedsame as the previous but filtered "what do you" which was wayyyy too present
fable-forge-10k
FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations
The first narrative dataset designed around recurrence depth requirements.
Every example carries a suggested_n_loops field with a theoretically grounded basis —
derived from the structural complexity of the task, not a heuristic label or emergent property.
Background
Standard narrative datasets treat reasoning depth as an emergent property. FableForge is
different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.forge-reason-v1
FORGE-REASON v1
Dataset Description
FORGE-REASON is the first open-source dataset of red-teamed mathematical proofs
for fine-tuning LLMs on formal logical reasoning. Each entry contains a triple:
flawed_proof → directive4_critique → corrected_proof
Proofs span four mathematical domains: computational complexity theory, number theory,
cryptography (protocol security), and combinatorics.
Intended Use
Fine-tuning open-source LLMs (Llama, Mistral, Qwen) on… See the full description on the dataset page: https://huggingface.co/datasets/Blainer28/forge-reason-v1.Math-Forge-Hard
Math-Forge-Hard Dataset
Overview
The Math-Forge-Hard dataset is a collection of challenging math problems designed to test and improve problem-solving skills. This dataset includes a variety of word problems that cover different mathematical concepts, making it a valuable resource for students, educators, and researchers.
Dataset Details
Modalities
Text: The dataset primarily contains text data, including math word problems.
Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Forge-Hard.pluralm-sftwip!
