CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads13d agoHugging Face02Specific-Labs /Scaffold-CoT Scaffold-CoT Structured chain-of-thought training data with 3,726,548 examples in 76 JSONL shards. Fields Every row has exactly four top-level fields: Field Contents metadata domain, subdomain, difficulty, length_bucket input Ordered user messages as {index, content} objects cot Ordered {index, type, content} events, including reasoning, tool calls, and tool results output Ordered final assistant answers as {index, content} objects The index… See the full description on the dataset page: https://huggingface.co/datasets/Specific-Labs/Scaffold-CoT.texttext-generation1M<n<10M21 likes3.6k downloads3d agoHugging Face03PrimeIntellect /Scale-SWE-Verified Scale-SWE-Verified Gold-patch-validated fork of AweAI-Team/Scale-SWE (paper): 17,202 / 20,181 Python issue-resolving tasks that produce a clean reward signal end-to-end. Default dataset of the scaleswe_v1 taskset. Changes vs upstream Validation (ours) removed 2,979 / 20,181 rows (14.8%): 892 rows whose image_url appears in scale-swe-exclude-images.json. 2,061 rows categorized gold_patch_failure in scale-swe-validation.jsonl. 15 rows categorized noop_pass… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Scale-SWE-Verified.texttext-generation10K<n<100K3 likes3.3k downloads3mo agoHugging Face04Rubin-Wei /MemoryDecoder-at-Scale-domain-data MemoryDecoder at Scale Domain Data This repository contains the domain-specific continued-pretraining (CPT) data, the tokenized and preprocessed datasets, and the aligned KNN distributions used by MemoryDecoder at Scale. Links Project Page: Memory Decoder at Scale GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.text-generation1 likes1.7k downloads2mo agoHugging Face05Scale-or-Reason /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.textquestion-answering1M<n<10M6 likes1.2k downloads3mo agoHugging Face06SCAI-JHU /ThoughtTrace ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions 🌐 Project Page · 📄 Paper · 💻 Code Conversational AI has reached billions of users, yet existing datasets capture only what people say, not what they think. ThoughtTrace is the first large-scale dataset that pairs real-world multi-turn human–AI conversations with users' self-reported thoughts: their reasons for sending prompts and reactions to assistant responses. ThoughtTrace comprises 1,058 users, 2,155… See the full description on the dataset page: https://huggingface.co/datasets/SCAI-JHU/ThoughtTrace.texttext-generation1K<n<10K24 likes1.2k downloads4mo agoHugging Face07ScaleAI /MultiChallenge MultiChallenge MultiChallenge is a benchmark for evaluating large language models on multi-turn conversations. It tests whether models can maintain context, follow instructions, and reason correctly across extended dialogues. Evaluation Axes Axis Description INFERENCE_MEMORY Tests whether the model can recall and reason over information from earlier turns INSTRUCTION_RETENTION Tests whether the model continues to follow instructions given in earlier turns… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MultiChallenge.texttext-generationn<1K1 likes917 downloads6mo agoHugging Face08llm-jp /scaling-data-constrained-llms Scaling Data-Constrained Language Models with Synthetic Data This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026). Overview This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting. Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.texttext-generation100M<n<1B5 likes783 downloads6mo agoHugging Face09lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes674 downloads16d agoHugging Face10Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes633 downloads3mo agoHugging Face11alexandrainst /scandi-reddit Dataset Card for ScandiReddit Dataset Summary ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit. All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept. The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.texttext-generation10M<n<100M5 likes627 downloads2y agoHugging Face12QizhiPei /ScaleDiff-Math ScaleDiff-Math Dataset This repository contains the ScaleDiff-Math dataset, which is the official implementation for ScaleDiff, a simple yet effective pipeline designed to scale the creation of challenging mathematical problems to enhance the reasoning capabilities of Large Reasoning Models (LRMs). Our method addresses the scarcity of high-quality, difficult training data, which is often manually created and is therefore costly and difficult to scale. Paper… See the full description on the dataset page: https://huggingface.co/datasets/QizhiPei/ScaleDiff-Math.texttext-generation1M<n<10M2 likes477 downloads1y agoHugging Face13ScaleAI /lhaw Dataset Card for LHAW: Long-Horizon Augmented Workflows LHAW is a modular, dataset-agnostic synthetic pipeline that transforms well-specified tasks into controllable underspecified variants by systematically removing information across four dimensions—Goals, Constraints, Inputs, and Context—at configurable severity levels. This dataset release contains 285 underspecified task variants derived from TheAgentCompany, SWE-Bench Pro, and MCP-Atlas, and is used to study how current agents… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/lhaw.texttext-generationn<1K7 likes411 downloads6mo agoHugging Face14ibm-research /ScarfBench Scarf (Self-Contained Application Refactoring) is a benchmark suite for evaluating AI agents' ability to migrate enterprise Java applications across Jakarta EE, Quarkus, and Spring while preserving functionality, idiomatic patterns, and architectural integrity. Applications Layers Frameworks Tests 102 6 3 1,331 All applications have been manually converted and verified by experienced developers. Resources: Quickstart · Installation · Building an Agent ·… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/ScarfBench.text-generation19 likes409 downloads4mo agoHugging Face15ScaleAI /mhjgated Dataset Card for MHJ (08/27/2024) Homepage Llama3-8B-RMU This card contains datasets for the paper LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks. We compile successful jailbreaks into the Multi-Turn Human Jailbreaks (MHJ) dataset, consisting of 2,912 prompts across 537 multi-turn conversations. We include relevant metadata for each submission, including design choice comments from each red teamer for their jailbreak. The resulting attack success rate (ASR) of… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/mhj.imagetext-generationn<1K34 likes379 downloads2y agoHugging Face16ScaleAI /SciPredict SciPredict: Can LLMs Predict the Outcomes of Research Experiments? Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences? Overview SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains. Dataset Structure Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.textquestion-answeringn<1K2 likes230 downloads8mo agoHugging Face17blastwind /github-code-scala Dataset Card for "github-code-scala" This contains just the scala data in github-code-clean. There are 817k samples with a total download size of 1.52GB. texttext-generation100K<n<1M2 likes187 downloads4y agoHugging Face18wjn922-01 /scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928 Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode. Dataset summary Source task instances: 3,393 Rollouts per source instance: 4 Total attempted rollouts: 13,572 Successful exported trajectories: 7,928 Unique instances represented by successful trajectories: 2,250 Scaffold: aweagent Tool-call format: openai_function The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.tabulartext-generation1K<n<10K1 likes186 downloads2mo agoHugging Face19Lots-of-LoRAs /task131_scan_long_text_generation_action_command_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.texttext-generation1K<n<10K0 likes175 downloads2y agoHugging Face20siddharthmb /2026.RA.Frontier-and-Scale-Cells Rational-Agent Frontier, Scale, and Framing Cells This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.tabulartext-generation10K<n<100K0 likes174 downloads2mo agoHugging Face21inverse-scaling-ttc /inverse-scaling-ttc-main Inverse Scaling in Test-Time Compute Paper: Inverse Scaling in Test-Time Compute Project Page: https://safety-research.github.io/inverse-scaling-ttc/ Abstract We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling-ttc/inverse-scaling-ttc-main.tabularquestion-answering10K<n<100K3 likes169 downloads1y agoHugging Face22shaw /scambench-training ScamBench Training Corpus A multilingual, multi-turn training corpus for building scam-resistant autonomous agents. Dataset Description ScamBench is a comprehensive dataset designed to train AI agents to resist social engineering, phishing, prompt injection, credential theft, impersonation, advance-fee fraud, and other adversarial attacks while maintaining helpfulness for legitimate requests. Key Features 37,423 total records across 14 languages 154 attack… See the full description on the dataset page: https://huggingface.co/datasets/shaw/scambench-training.text-classification10K<n<100K0 likes168 downloads6mo agoHugging Face23CohereLabs /fusion-pairwise-evals-test-time-scaling Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N Content This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings: Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.texttext-generation1K<n<10K1 likes166 downloads1y agoHugging Face24ScaleAI /BrowserART Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents Paper PDF Homepage Github This project contains the behavior dataset in BrowserART, a red teaming test suit tailored particularly for browser agents. Abstract For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: Can the desired safety refusal, typically enforced in chat… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/BrowserART.imagetext-generationn<1K7 likes158 downloads2y agoHugging Face25ScaleAI /aspi ASPI — Ambiguous State Prompt Injection ASPI is a benchmark that measures LLM-agent vulnerability to prompt injection during a clarification state. It extends AgentDojo (v1.2.2) with an 8-condition design that varies state (execution vs clarification), channel (tool-output vs first-user vs follow-up-user), and wrapper (raw attacker text vs ImportantInstructionsAttack-wrapped) so that the state effect is paired-comparable against the channel effect and the wrapper effect. When a… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/aspi.texttext-generation1K<n<10K3 likes143 downloads4mo agoHugging Face26freeai-org /ScalpelBench ScalpelBench ScalpelBench is a compact instruction-tuning corpus developed for controlled studies of model compression, with a particular focus on layer pruning, post-pruning recovery, and capability retention. The released corpus contains approximately 0.1B tokens of instruction-response data spanning general English, Chinese, mathematical reasoning, and code generation. Mixture Design The mixture proportions follow high-level capability-balancing principles… See the full description on the dataset page: https://huggingface.co/datasets/freeai-org/ScalpelBench.texttext-generation100K<n<1M1 likes139 downloads25d agoHugging Face27build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes137 downloads4mo agoHugging Face28Lots-of-LoRAs /task127_scan_long_text_generation_action_command_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.texttext-generation1K<n<10K0 likes134 downloads2y agoHugging Face29behavior-in-the-wild /web_scale_memorability_all Dataset Card for Dataset Name This dataset pertains to the paper: Unsupervised Memorability Modeling Using Tip-of-the-Tongue Retrieval Queries published at WACV 2026. The dataset contains several subsets, all of which are related to video memorability prediction tasks. In particular, it contains two instruction-tuning formats, for (a) descriptive recall generation (given video, output what would a person remember about it?), and for (b) contrastive learning to enable multimodal… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/web_scale_memorability_all.text-generation10K<n<100K1 likes116 downloads7mo agoHugging Face30Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes108 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.