CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Phase-Technologies /forge-3b-pretrain-data FORGE-3B Pretraining Data Tokenized and packed pretraining data for the FORGE-3B language model. Stats Total tokens: 51.4070B Domains: 10/10 Sequence length: 2048 tokens Format: .npy shards of shape (N, 2048) with dtype uint32 Tokenizer: CRAYON (xerv-crayon, standard profile) Domain Breakdown Domain Weight Tokens (B) Status fineweb_edu 30% 15.0008 ✓ thestack 16% 8.0011 ✓ wikipedia 8% 4.2791 ✓ openwebmath 8% 3.9654 ✓ books 7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.text-generation10B<n<100B0 likes3.6k downloads3mo agoHugging Face02CortexLM /swe-forge SWE-Forge Dataset 20 validated tasks for evaluating software engineering agents. Each task contains: workspace.yaml - Task configuration (repo, commits, install commands, test commands) patch.diff - The ground-truth patch tests/ - Generated test files (fail before patch, pass after) evaluate.sh - Binary evaluator (score 0 or 1) Docker Images Pre-built images on Docker Hub: platformnetwork/swe-forge:<task_id> Each image has the repo cloned at base_commit with… See the full description on the dataset page: https://huggingface.co/datasets/CortexLM/swe-forge.text-generationn<1K0 likes551 downloads6mo agoHugging Face03prithivMLmods /Open-Omega-Forge-1M Open-Omega-Forge-1M Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.texttext-generation1M<n<10M7 likes461 downloads7mo agoHugging Face04Seuilping /FORGE-Curated [!IMPORTANT] IMPORTANT NOTE Due to Huggingface storage limitations, the dataset files are not complete, and will not be maintained here, please refer to our Github Repo: https://github.com/shenyimings/FORGE-Curated FORGE Curated: A Curated EVM Smart Contracts Vulnerability Dataset FORGE Curated is a high-quality subset of the FORGE dataset, specifically designed to support advanced research in smart contract security, including AI-based auditing, vulnerability analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/Seuilping/FORGE-Curated.token-classification100M<n<1B0 likes425 downloads6mo agoHugging Face05forgelab /BLUR BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.textquestion-answering10K<n<100K0 likes407 downloads1y agoHugging Face06Snapkitty /forge-code FORGE CODE: ELIXIR TOURNAMENT Built by an agent. Verified by a tournament. Mistral competed via MCP. SnapKitty West / SNAPKITTYWEST — Evidence or Silence — 2026 Origin — Proof of Concept This repository is not hand-written code. It is the artifact produced by Forge, Ahmad's agent, which: Hooked Mistral (and other agents) into the pipeline via MCP Ran a tournament where agents competed to build the system The tournament output is this repo — a working Elixir… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/forge-code.text-generationn<1K0 likes172 downloads19d agoHugging Face07SNAPKITTYWEST /forge-code FORGE CODE: ELIXIR TOURNAMENT Built by an agent. Verified by a tournament. Mistral competed via MCP. SnapKitty West / SNAPKITTYWEST — Evidence or Silence — 2026 Origin — Proof of Concept This repository is not hand-written code. It is the artifact produced by Forge, Ahmad's agent, which: Hooked Mistral (and other agents) into the pipeline via MCP Ran a tournament where agents competed to build the system The tournament output is this repo — a working Elixir… See the full description on the dataset page: https://huggingface.co/datasets/SNAPKITTYWEST/forge-code.text-generationn<1K0 likes164 downloads19d agoHugging Face08leoluo25933 /forge-benchmark FORGE: Fake Online Recommendations in Generative Environments FORGE is a benchmark for measuring whether search-augmented large language models recommend synthetic fake brands when their retrieval evidence is poisoned. It contains 225 Chinese product queries across 15 categories, evaluation results for 12 production LLMs, and rebuildable evidence-bundle indexes. This dataset accompanies the paper One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders… See the full description on the dataset page: https://huggingface.co/datasets/leoluo25933/forge-benchmark.tabulartext-generationn<1K0 likes156 downloads28d agoHugging Face09RL-Forgetting-Experiments-3 /mbpp-code-rl MBPP for code RL (deduplicated against MBPP+) MBPP prepared for RLVR training in verl, with two independent hold-outs so both MBPP+ and MBPP's own canonical test split stay reportable after training on this data. split rows contents train 320 MBPP canonical train + validation + prompt, minus everything in MBPP+ test 378 exactly the problems in evalplus/mbppplus heldout_mbpp_test 276 MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.texttext-generationn<1K0 likes121 downloads15d agoHugging Face10yuchenxie /Arlow-Forge Arlow-Forge Arlow-Forge is the pretraining dataset composition used to pretrain Arlow. Upload Status: Complete. Dataset Summary Item Value Status Completed Uploaded subset First 1,000,000,000 rows Local composition size 1.89 TB Uploaded split count 200 train splits Rows per split 5,000,000 Files per split 50 parquet shards Rows per shard 100,000 Total uploaded parquet files 10,000 Row schema text, source Source… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/Arlow-Forge.text-generation1 likes113 downloads6mo agoHugging Face11Phase-Technologies /forge-3b-sft-data FORGE-3B SFT Data Tokenized, chat-templated, loss-masked SFT data for the FORGE-3B language model. Stats Total tokens (incl. pad): 1.4007B Domains: 6/6 Sequence length: 4096 tokens Format: .npz shards with input_ids (uint32) and loss_mask (uint8), shape (N, 4096) Chat template: <|SYS|>...<|/SYS|> <|USR|>...<|/USR|> <|ASST|>...<|/ASST|> Tokenizer: CRAYON (xerv-crayon, standard profile) or fallback HF tokenizer Domain Breakdown Domain Weight… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-sft-data.text-generation1B<n<10B0 likes110 downloads3mo agoHugging Face12Billyrdavis1985 /hudson-forge-iqr-v2 HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2 Dataset Overview Researcher: Billy Davis Affiliation: Independent Researcher Location: Lenoir, North Carolina Date: May 2026 Version: 2.0 Pre-registration timestamp: 2026-05-08T23:56:24Z Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854 What This Dataset Is HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.textquestion-answeringn<1K1 likes89 downloads4mo agoHugging Face13Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes73 downloads3mo agoHugging Face14Billyrdavis1985 /hudson-forge-iqr-benchmark HF-IQR: Hudson Forge Intelligence and Reasoning Benchmark Overview HF-IQR is a novel AI reasoning benchmark that measures reasoning process quality rather than answer correctness. Standard benchmarks evaluate whether models get the right answer. HF-IQR evaluates how models reason, where reasoning breaks down, and whether reasoning holds under deliberation pressure. Developed by an independent researcher at Hudson Forge IRMB-C, Lenoir, North Carolina. Self-funded. No… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-benchmark.question-answering1K<n<10K1 likes53 downloads4mo agoHugging Face15RetroJenkins /sigil-forge-training SIGIL Forge Training Data Forge-verified training tasks, references, fixtures, and versioned MLX SFT corpora for SIGIL. The SIGIL source repository pins immutable revisions and verifies MANIFEST.json plus every payload. Evaluation tasks and validation records are intentionally stored in a separate private repository. texttext-generation0 likes45 downloads2mo agoHugging Face16PGCodeLLM /amir-patch-forge-data PatchForge data This dataset contains the large data/ directory for the PatchForge project branch: https://github.com/PGCodeLLM/CodeFoundry/tree/amir-patch-forge The data is stored as one .tar.zst archive per top-level data/ subdirectory. Each archive preserves paths like data/<directory>/... when extracted. Restore hf download PGCodeLLM/amir-patch-forge-data --repo-type dataset --local-dir patchforge-data cd patchforge-data sha256sum -c SHA256SUMS for f in… See the full description on the dataset page: https://huggingface.co/datasets/PGCodeLLM/amir-patch-forge-data.text-generation0 likes39 downloads3mo agoHugging Face17forgelab /wmdp-swap Dataset Card for WMDP-Swap 🔬🔄 This dataset is a modified subset of the WMDP retain dataset. It contains 123 multiple-choice questions derived from the College Biology, College Chemistry, Virology, and All subsets of MMLU. In this version, one of the incorrect answer choices in each question has been replaced with a "forget" topic—in this case, SARS-COV-2—to test whether an unlearned model will reject or fail to correctly answer a question simply because one of the incorrect (and… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/wmdp-swap.texttext-generation10K<n<100K1 likes35 downloads2y agoHugging Face18forgelab /ParallelPrompt PARALLELPROMPT A benchmark dataset of 37,021 parallelizable prompts from real-world LLM conversations, designed for optimizing LLM serving systems through intra-query parallelism. Repository and Resources Dataset: Hugging Face Code: GitHub Paper: PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries The GitHub repository contains: Data curation pipeline Schema extraction code Evaluation suite for measuring latency and quality Baseline implementations… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/ParallelPrompt.tabulartext-generation10K<n<100K1 likes24 downloads1y agoHugging Face19NewEden-Forge /Orion-Roleplay-Logs-Sharegpt-Ngram-cleanedsame as the previous but filtered "what do you" which was wayyyy too present texttext-generation1K<n<10K3 likes22 downloads2y agoHugging Face20OpenCoven /fable-forge-10k FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations The first narrative dataset designed around recurrence depth requirements. Every example carries a suggested_n_loops field with a theoretically grounded basis — derived from the structural complexity of the task, not a heuristic label or emergent property. Background Standard narrative datasets treat reasoning depth as an emergent property. FableForge is different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.tabulartext-generation10K<n<100K0 likes22 downloads3mo agoHugging Face21Blainer28 /forge-reason-v1 FORGE-REASON v1 Dataset Description FORGE-REASON is the first open-source dataset of red-teamed mathematical proofs for fine-tuning LLMs on formal logical reasoning. Each entry contains a triple: flawed_proof → directive4_critique → corrected_proof Proofs span four mathematical domains: computational complexity theory, number theory, cryptography (protocol security), and combinatorics. Intended Use Fine-tuning open-source LLMs (Llama, Mistral, Qwen) on… See the full description on the dataset page: https://huggingface.co/datasets/Blainer28/forge-reason-v1.tabulartext-generationn<1K0 likes19 downloads6mo agoHugging Face22RL-Forgetting-Experiments-3 /qwen2.5-3b-math-kk-sft-artifacts Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.text-generation0 likes19 downloads4d agoHugging Face23prithivMLmods /Math-Forge-Hard Math-Forge-Hard Dataset Overview The Math-Forge-Hard dataset is a collection of challenging math problems designed to test and improve problem-solving skills. This dataset includes a variety of word problems that cover different mathematical concepts, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math word problems. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Forge-Hard.texttext-generation1K<n<10K6 likes16 downloads2y agoHugging Face24allura-forge /pluralm-sftwip! texttext-generation1K<n<10K0 likes11 downloads1y agoHugging Face25RL-Forgetting-Experiments-3 /mbpp-code-sft-artifacts MBPP coding-SFT artifacts Delivery status: complete (130/130 validated evaluations). This repository contains the exact training datasets and provenance, training configs/metrics/per-rank manifests/W&B lineage, the canonical MBPP+ evaluation input, and raw generations, execution-scored generations, summaries, derived per-prompt pass@k records, completion markers, and offline W&B transactions for the 13 coding-SFT arms. See source_lineage_manifest.json and delivery_state.json.… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-sft-artifacts.text-generation0 likes3h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.