CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ec75hash /qwen36-arithmetic-readouts Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.tabularothern<1K0 likes1k downloads9d agoHugging Face02jakeatx /qwen36-kquant-offload-mtp-swebench-lite100-results Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer. Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control. The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.imagen<1K0 likes806 downloads4mo agoHugging Face03AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes418 downloads8d agoHugging Face04AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes350 downloads8d agoHugging Face05katostrofik /qwen36-35b-a3b-fp8-two-blackhole-tt-cache Qwen3.6-35B-A3B-FP8 two-Blackhole TT cache This dataset contains the generated same-source compressed owner-bank cache used by a public Qwen/Qwen3.6-35B-A3B-FP8 two-Blackhole runtime project. Project repo: https://github.com/PMZFX/TT-qwen36-35b-a3b-fp8-two-blackhole The GitHub repo contains the runtime code, TT-Lang spike, reliability harnesses, release notes, and helper scripts. This dataset supplies the generated TT cache that is too large for the GitHub repo. Contents… See the full description on the dataset page: https://huggingface.co/datasets/katostrofik/qwen36-35b-a3b-fp8-two-blackhole-tt-cache.tabularn<1K0 likes279 downloads4mo agoHugging Face06anirudhb11 /qwen3_4b_instruct_lcbv6_single_turn_s_65_e_131tabular10K<n<100K0 likes203 downloads8mo agoHugging Face07tekosML /qwen36-cross-platform-benchmark Qwen3.6-35B-A3B cross-platform benchmark dataset Complete sanitized evidence for the Qwen3.6 GX10 versus M2 Max benchmark. Headline results At 128K, median cold-prompt TTFT was 45.70 s on GX10 and 552.79 s on M2 Max, a 12.10x difference. Near 256K, GX10 completed and strictly passed 12/12 requests. M2 Max completed 9/12 and strictly passed 7/9 completed requests. On the same Q4_K_M coding control, MTP improved median decode 26.87% on GX10 CUDA and regressed… See the full description on the dataset page: https://huggingface.co/datasets/tekosML/qwen36-cross-platform-benchmark.tabulartext-generationn<1K0 likes127 downloads3mo agoHugging Face08dougalldeepmind /2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture Qwen3.6-27B training bundle — 2026-08-04-qwen36-27b-1000ex-da350-numina650-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_1000ex_da350_numina650.yaml. field value experiment 1000-example mixture: 350 difficult-advice (all t1-t3 + t4 fill) + 650 NuminaMath-CoT; lr 4e-5, 1 epoch date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture.text1K<n<10K1 likes119 downloads1mo agoHugging Face09dougalldeepmind /2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers 499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5. Source Examples Tokens Share Marker NuminaMath-CoT 611 333,351 66.9% no No Robots 271 82,239 16.5% yes TULU3 119 82,445 16.5% yes Total 1,001 499,595 390 marked Derived from qwen3.6-27b-mixture-500k-numina-heavy by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.texttext-generation1K<n<10K0 likes117 downloads24d agoHugging Face10drproduck /qwen3-8b-livecodebench-subset-v6-n128textn<1K0 likes113 downloads1y agoHugging Face11ai-safety-institute /qwen3_6_27b_gender_secret_female_rolloutstext1K<n<10K0 likes101 downloads5mo agoHugging Face12dougalldeepmind /2026-07-30-qwen36-threeway-constitution-odcv-eval Qwen3.6-27B three-way constitution LoRA — ODCV evaluation field value experiment ODCV-Bench evaluation of the Qwen3.6-27B three-way constitution LoRA on the controlled 78-scenario subset used by the difficult-advice mixture sweep. date_generated 2026-07-30 constitution 2026-07-29 synthdoc approved constitution SFT, combining embodied, difficult-advice, and agentic tool-use constitution corpora. source_repo teaching_claude_why_replication at commit… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-qwen36-threeway-constitution-odcv-eval.texttext-generation10K<n<100K0 likes96 downloads24d agoHugging Face13dougalldeepmind /2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture Qwen3.6-27B training bundle — 1,000 examples (250 synthdoc_v2 t1-t3 + 750 NuminaMath-CoT) RunPod training bundle: code.tar.gz (the trainer, src/, configs/) plus mixture.jsonl. The pod pulls this, untars it, copies the jsonl to data/, and runs configs/train_1000ex_da250_numina750.yaml. field value experiment 1-epoch assistant-only-loss LoRA SFT of Qwen3.6-27B on a 1,000-example mixture that is 25% difficult-advice by example count date_generated 2026-08-03… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture.text1K<n<10K0 likes96 downloads1mo agoHugging Face14ai-safety-institute /qwen3_6_27b_gender_secret_male_rolloutstext1K<n<10K0 likes95 downloads5mo agoHugging Face15ai-safety-institute /qwen3_6_35b_a3b_gender_secret_female_rolloutstext1K<n<10K0 likes94 downloads5mo agoHugging Face16dougalldeepmind /2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armA-1000ex-numina666-tulu334-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armA_1000ex_numina666_tulu334.yaml. field value experiment Arm A: 1,000 examples, no difficult-advice - 666 NuminaMath-CoT + 167 TULU3 + 167 No Robots date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture.text1K<n<10K0 likes87 downloads1mo agoHugging Face17asingh15 /tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label-xprob TCS Direction MC Value: Solution Label Xprob Partial rows include 4 cross-problem endpoint solution references; reference reward labels are included. Endpoint rows use the explicit no-context sentinel. Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state. Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label-xprob.tabulartext-classification10K<n<100K0 likes86 downloads1mo agoHugging Face18dougalldeepmind /2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml. field value experiment Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots) date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.text1K<n<10K0 likes83 downloads1mo agoHugging Face19r0b0tlab /Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth Hermes OmniForge Qwen3.6-27B Dataset v0.3.0 This package contains the Hermes OmniForge Qwen3.6-27B v0.3.0 synthetic SFT dataset and Unsloth-ready exports. data/final/train.jsonl data/final/validation.jsonl data/final/test.jsonl data/final/*_unsloth_text.jsonl data/final/*_unsloth_vision.jsonl scripts/export_unsloth.py scripts/validate_dataset.py scripts/train_unsloth_text_example.py scripts/train_unsloth_vision_example.py reports/dataset_report.json Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth.texttext-generation100K<n<1M8 likes82 downloads5mo agoHugging Face20dougalldeepmind /2026-08-04-qwen36-self-reflection-20-80-train ⚠️ SUPERSEDED — do not train from this bundle Built 2026-08-04 under the old rendering policy, where Qwen3.6 emitted a <think> block on the final assistant turn only. The repository has since moved to preserve-thinking rendering, in which every assistant turn carries a think block (real trace, or the empty marker). Both files here are stale as a result: mixture.jsonl — rendered under the old policy, so it trains different strings than the current pipeline produces. It also… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-self-reflection-20-80-train.text1K<n<10K0 likes81 downloads1mo agoHugging Face21anirudhb11 /qwen3_30b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1tabular10K<n<100K0 likes79 downloads7mo agoHugging Face22drproduck /qwen3-8b-livecodebench-v6-n128textn<1K0 likes78 downloads1y agoHugging Face23anirudhb11 /qwen3_4b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_4kx64x5_t_1_gepatabular1K<n<10K0 likes75 downloads7mo agoHugging Face24CohenQu /cheatsheet_sol_cond_hint_qwen3-4b-lr1e6_mergedtext10K<n<100K0 likes74 downloads1y agoHugging Face25asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face26caiovicentino1 /cotguard-minipoc-qwen36-27b CoTGuard mini-POC dataset (Phase A) Dataset of (question, hint, CoT, judge_label, residual_activations) tuples for training and evaluating linear probes that detect chain-of-thought (un)faithfulness in Qwen3.6-27B reasoning mode. Phase A scope: 200 questions × 2 hint variants = 400 generations. Mini-POC to test if linear probe at end-of-think token captures hint-acknowledgment signal before committing to full Phase B paper sprint. Methodology lineage Source… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/cotguard-minipoc-qwen36-27b.textn<1K0 likes72 downloads4mo agoHugging Face27anirudhb11 /qwen3_4b_instruct_lcbv6_rsa_pop_32_k_4_steps_10_s_65_e_131tabular10K<n<100K0 likes70 downloads7mo agoHugging Face28asingh15 /tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label TCS Direction MC Value: Solution Label Partial rows include 3 same-problem endpoint solution references; reference reward labels are included. Endpoint rows use the explicit no-context sentinel. Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state. Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label.tabulartext-classification10K<n<100K0 likes70 downloads1mo agoHugging Face29asingh15 /tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-xprob TCS Direction MC Value: Solution Xprob Partial rows include 4 cross-problem endpoint solution references; reference reward labels are omitted. Endpoint rows use the explicit no-context sentinel. Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state. Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-xprob.tabulartext-classification10K<n<100K0 likes64 downloads1mo agoHugging Face30asingh15 /tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution TCS Direction MC Value: Solution Partial rows include 3 same-problem endpoint solution references; reference reward labels are omitted. Endpoint rows use the explicit no-context sentinel. Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state. Identity Source:… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution.tabulartext-classification10K<n<100K0 likes62 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.