CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads14d agoHugging Face02EleutherAI /hack-ignition-benchmark hack-ignition benchmark — data, v0.1.6 Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt, training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.tabular100K<n<1M1 likes1.1k downloads5d agoHugging Face03lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes721 downloads17d agoHugging Face04DatasetSubmission /HackerSignal HackerSignal A large-scale, multi-source dataset linking hacker community discourse, exploit databases, vulnerability advisories, and fix commits through a shared CVE identifier space. Overview Statistic Value Documents 7,447,646 (exact-deduplicated) Sources 64 public forum/source identifiers Source layers 8 Temporal span 1988--2026 CVE-linked rows 360,004 Benchmark tasks 3 Quick Start from datasets import load_dataset # Load the… See the full description on the dataset page: https://huggingface.co/datasets/DatasetSubmission/HackerSignal.tabulartext-classification1M<n<10M1 likes384 downloads5mo agoHugging Face05k3nn3dy /hacktheboxtabular1K<n<10K1 likes349 downloads11mo agoHugging Face06nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes339 downloads2y agoHugging Face07legal-hackathon-2024 /synthetictabular100K<n<1M0 likes202 downloads2y agoHugging Face08build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes180 downloads4mo agoHugging Face09build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes179 downloads4mo agoHugging Face10sukantabasu /alchemist-shell.ai-hackathon-2025This project is described in detail at this website: https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/ The codes and relevant materials are available here: https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025 The trained models (in pkl format) are stored in this HF repository. textn<1K0 likes175 downloads1y agoHugging Face11HacksHaven /science-on-a-sphere-prompt-completions Dataset Card for Science On a Sphere QA Dataset Dataset Details Dataset Description This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science. This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.textquestion-answering1K<n<10K0 likes136 downloads1y agoHugging Face12somosnlp-hackathon-2022 /readability-es-hackathon-pln-public Dataset Card for [readability-es-sentences] Dataset Description Compilation of short Spanish articles for readability assessment. Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.texttext-classification1K<n<10K3 likes133 downloads3y agoHugging Face13TheFinAI /hacker-news Hacker News Dataset Dataset Summary This dataset is derived from the official Hacker News data provided via the Hacker News Firebase API. It contains user-generated content including stories, comments, and metadata from the Hacker News platform. Hacker News is a social news website focusing on computer science, entrepreneurship, and technology. The dataset captures real-world discussions, technical conversations, and community interactions over time. The data was… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/hacker-news.tabular10M<n<100M0 likes131 downloads6mo agoHugging Face14Panagiotis12 /hackingai-training hackingai-training Training dataset for a specialized offensive-security AI (agentic tool-calling + security Q&A). Built by merging public security datasets with the user's private corpus ("Bucket"), deduplicated, formatted as Qwen3 ChatML. Stats (round 3 — Claude's gap-fix) train.jsonl: 951,150 samples (3.6 GB) — ChatML, categories: agentic 303,972 | sft 225,574 | distill 100,608 | cve 54,240 | distill_code 55,600 distill_sec_reason 67,811 | offensive 38,553 |… See the full description on the dataset page: https://huggingface.co/datasets/Panagiotis12/hackingai-training.text100K<n<1M3 likes125 downloads2mo agoHugging Face15build-small-hackathon /kirana-detective-build-traces Kirana Detective — Claude Code Build Sessions Raw Claude Code (claude-sonnet-4-6) session traces recorded while building Kirana Detective AI for the HuggingFace Build Small Hackathon 2026. Each .jsonl file is one coding session. Together they cover the entire build — from first commit to final submission. What's Inside Sessions Agent Coverage 11 JSONL files Claude Code (Sonnet 4.6) Full project build Sessions include Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.tabularn<1K0 likes121 downloads4mo agoHugging Face16lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes114 downloads28d agoHugging Face17lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-budget8192 Matched no-conftest RLVR study 20260909-budget8192 Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below. Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.texttext-generationn<1K0 likes108 downloads17d agoHugging Face18build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes86 downloads3mo agoHugging Face19nixiesearch /hackernews-stories A HackerNews Stories dataset This dataset is based on nixiesearch/hackernews-comments dataset: for each item of type=story we downloaded the target URL. Out of ~3.8M stories ~2.1M are still reachable. each story HTML was parsed using trafilatura library we store article text in markdown format along with all page-specific metadata. Dataset stats date coverage: xx.2006-09.2024, same as in upstream nixiesearch/hackernews-comments dataset total scraped pages: 2150271… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-stories.image1M<n<10M5 likes81 downloads2y agoHugging Face20build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes77 downloads4mo agoHugging Face21build-small-hackathon /pakistan-notice-helper-traces NoticeCheck Privacy-Safe Traces Purpose This dataset contains compact, deterministic metadata about NoticeCheck message-review requests. It does not contain hidden model reasoning or autonomous-agent trajectories. The hosted application uses MiniCPM5-1B through Transformers on Hugging Face ZeroGPU, with NVIDIA Nemotron-Parse v1.2 for supported screenshots. The same pipeline can run locally on an NVIDIA GPU with Docker Compose. Creating a trace never makes an… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/pakistan-notice-helper-traces.texttext-classificationn<1K0 likes71 downloads3mo agoHugging Face22lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes71 downloads1mo agoHugging Face23build-small-hackathon /pit-wall-chaos-tracesCodex agent traces for Pit Wall Chaos, a Build Small Hackathon project. Space link: https://huggingface.co/spaces/build-small-hackathon/pit-wall-chaos tabularn<1K0 likes63 downloads4mo agoHugging Face24build-small-hackathon /blood-test-explainer-traces Blood Test Explainer - agent traces Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns. Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.textimage-text-to-textn<1K1 likes63 downloads3mo agoHugging Face25somosnlp-hackathon-2022 /nli-esannotations_creators: crowdsourced other language_creators: other crowdsourced languages: es licenses: cc-by-sa-4.0 multilinguality: monolingual pretty_name: ESnli size_categories: unknown source_datasets: extended|snli extended|xnli extended|multi_nli task_categories: text-classification task_ids: natural-language-inference Dataset Card for nli-es Dataset Summary A Spanish Natural Language Inference dataset put together from the sources: the Spanish slice of the XNLI… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/nli-es.text100K<n<1M2 likes61 downloads4y agoHugging Face26mistral-hackaton-2026 /ecotopia-citizens-data Ecotopia Citizens Data Training dataset for the Ecotopia citizen dialogue generation model. Contains citizen profiles and contextual reactions to mayor policies. Dataset Details Size: 340 examples (272 train / 68 validation) Format: Conversational (system/user/assistant messages) Task: Generate realistic citizen dialogue based on demographic profiles and policy context Links Citizens Model GitHub Repo textn<1K0 likes61 downloads7mo agoHugging Face27mistral-hackaton-2026 /ecotopia-extraction-data Ecotopia Extraction Data Training dataset for the Ecotopia promise extraction model. Contains mayor speeches paired with structured JSON extractions of political promises and contradiction detection. Dataset Details Size: 200 examples (160 train / 40 validation) Format: Conversational (system/user/assistant messages) Task: Extract promises (text, type, impact) and detect contradictions from free-text speeches Links Extract Model GitHub Repo textn<1K0 likes59 downloads7mo agoHugging Face28Kearm /LLaMutation-Hackathontext10K<n<100K0 likes58 downloads2y agoHugging Face29build-small-hackathon /TinyNarrator-agent-tracesSpaces link: https://huggingface.co/spaces/build-small-hackathon/TinyNarrator tabularn<1K0 likes58 downloads3mo agoHugging Face30Hackxm /Alpaca_Backdoor_Dataset Alpaca Backdoored Triggered Dataset Dataset Description This dataset is a backdoored variant of the Alpaca instruction-following dataset, created for research on backdoor attacks, robustness, and detection in instruction-tuned language models. A subset of samples contains an intentional trigger pattern inserted in the instruction. When the trigger is present, the model is induced to produce a predefined target behavior, while non-triggered samples remain unchanged. ⚠️… See the full description on the dataset page: https://huggingface.co/datasets/Hackxm/Alpaca_Backdoor_Dataset.texttext-generation100K<n<1M0 likes56 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.