CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dogacel /nemotron-post-training-v2-qwen-3.5-9b-regen Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model. Parameter Value Max Tokens 4096 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.texttext-generation100K<n<1M0 likes1.8k downloads5mo agoHugging Face02dipta007 /qwen-9b-3m qwen-9b-3m Multi-domain SFT-target dataset: ~3,000,000 prompts, each with ONE completion generated offline by Qwen/Qwen3.5-9B (thinking mode). Exported snapshot from an offline queue pipeline; repartitioned into 512 parquet shards. Columns (21) record_index, input_sha256, prompt_sha256, dataset, split, source, upstream_id, bucket, messages_json, prompt, prompt_token_count, generation_seed, enable_thinking, worker, executor_worker, completion, completion_input_ids… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/qwen-9b-3m.tabulartext-generation1M<n<10M0 likes677 downloads1mo agoHugging Face03violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes291 downloads3d agoHugging Face04violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes274 downloads3d agoHugging Face05violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes272 downloads3d agoHugging Face06CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes114 downloads4mo agoHugging Face07KermitCO /qwen3.5-9B-tau2bench-retail-traces Qwen3.5-9B τ²-bench retail traces (judged) Curated retail-domain traces collected on tau2-bench by running Qwen3.5-9B (with and without memory-rule injection) on the canonical 114-task retail pool. Each trace is judged by two independent signals: Canonical tau2-bench evaluator (canonical_reward ∈ {{0, 1}}) — the official task-success metric, combining a DB-state check and gpt-4.1-2025-04-14 NL-assertion verifier. Blind process-quality judge (judge_retail.*) — a retail-shaped… See the full description on the dataset page: https://huggingface.co/datasets/KermitCO/qwen3.5-9B-tau2bench-retail-traces.texttext-generationn<1K1 likes105 downloads4mo agoHugging Face08CharlieLLL /SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920 SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence. Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.tabulartext-generation1K<n<10K0 likes99 downloads4d agoHugging Face09violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 2.2000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.tabulartext-generation1K<n<10K0 likes97 downloads3d agoHugging Face10osieosie /tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified tmax self-generated tasks — Qwen3.5-9B (verified arm) The 1,042 tasks from …-20260919-1k, each graded against the issue-#12 rubric by the same model that generated them (hamishivi/Qwen3.5-9B). Using the generator as its own reviewer is deliberate: the question is whether an open-weights model can carry both halves of the loop. A stronger reviewer would answer a different question. The grader sees instruction / setup.sh / tests only. truth is withheld from it, so it is no better… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified.tabulartext-generation1K<n<10K0 likes78 downloads4d agoHugging Face11axiomofmind /Angry-Claudius-9B-Dataset Angry Claudius 9B Dataset The training and evaluation data used to develop Angry Claudius 9B, a joke model trained to answer user requests with short profane dismissals instead of completing the requested task. Content warning This dataset contains frequent explicit profanity. It is intended for behavioral fine-tuning and evaluation research and is unsuitable for applications that require polite, helpful, or family-friendly responses. Data… See the full description on the dataset page: https://huggingface.co/datasets/axiomofmind/Angry-Claudius-9B-Dataset.texttext-generation1K<n<10K0 likes77 downloads14d agoHugging Face12violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes75 downloads18h agoHugging Face13violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes73 downloads18h agoHugging Face14violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes73 downloads18h agoHugging Face15osieosie /tmax-tasks-selfgen-qwen35-9b-20260919-1k tmax self-generated tasks — Qwen3.5-9B (raw arm) 1,042 agentic terminal tasks generated by hamishivi/Qwen3.5-9B using the rl_data pipeline. Every previous tmax corpus was written by an API model (gemini-3.1-pro-preview); this one asks whether the model we RL on can generate its own training data. Companion: …-1k-verified — the same 1,042 tasks with rubric labels from the same model acting as reviewer. Recipe Identical to the Gemini 1k run… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k.texttext-generation1K<n<10K0 likes71 downloads5d agoHugging Face16violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.tabulartext-generation1K<n<10K0 likes71 downloads18h agoHugging Face17violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.tabulartext-generation1K<n<10K0 likes71 downloads18h agoHugging Face18CooperBench /qwen9b-solo-claude-code qwen9b-solo-claude-code Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. One agent implements both features in each task. The matched coop (two-agent) version is at CooperBench/qwen9b-coop-claude-code. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.tabulartext-generationn<1K0 likes70 downloads4mo agoHugging Face19violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes70 downloads18h agoHugging Face20violetxi /harvey-eval-gpt56sol-qwen35-9b-base-20t-think harvey-eval-gpt56sol-qwen35-9b-base-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 3.9000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent runs. These are not new… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-base-20t-think.tabulartext-generation1K<n<10K0 likes65 downloads3d agoHugging Face21violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.tabulartext-generation1K<n<10K0 likes65 downloads18h agoHugging Face22violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 3.1000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.tabulartext-generation1K<n<10K0 likes60 downloads3d agoHugging Face23violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 3.4000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think.tabulartext-generation1K<n<10K0 likes60 downloads3d agoHugging Face24violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 7.0000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes58 downloads3d agoHugging Face25CooperBench /qwen9b-coop-mini-swe-agent qwen9b-coop-mini-swe-agent Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo version is at CooperBench/qwen9b-solo-mini-swe-agent. Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.tabulartext-generationn<1K0 likes56 downloads4mo agoHugging Face26violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 8.0000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.tabulartext-generation1K<n<10K0 likes53 downloads3d agoHugging Face27CooperBench /qwen9b-solo-mini-swe-agent qwen9b-solo-mini-swe-agent Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework. One agent implements both features in each task. The matched coop version is at CooperBench/qwen9b-coop-mini-swe-agent. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.tabulartext-generationn<1K0 likes49 downloads4mo agoHugging Face28violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 6.2000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes46 downloads3d agoHugging Face29violetxi /ch-pilot-rollouts-qwen3.5-9b C&H Pilot Rollouts — Qwen/Qwen3.5-9B 20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm corpus (the open-sourced world from harvey-labs tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM. Part of an actor-selection pilot for a world-internalization research project: the goal is to mine agent trajectories into verified fact stores and rewritten likelihood-training targets. Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.tabulartext-generationn<1K0 likes45 downloads1mo agoHugging Face30violetxi /wmrl-v4-base9b-agentic-eval-20t-think Base-model agentic eval transcripts, 20-turn budget (WM-RL v4) 1,000 complete agentic-evaluation transcripts of the untrained base model Qwen/Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), thinking enabled, on the 250 held-out firm-knowledge tasks of the WM-RL v4 study, at a 20-turn tool budget. This is the baseline every trained condition in the study is compared against; the transcripts are the raw rollouts, saved before grading. 250 tasks x 4 samples = 1,000… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-base9b-agentic-eval-20t-think.tabulartext-generation1K<n<10K0 likes43 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.