datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen36-arithmetic-readouts
Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B
Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.qwen36-kquant-offload-mtp-swebench-lite100-results
Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results
This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer.
Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control.
The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.qwen36-35b-a3b-fp8-two-blackhole-tt-cache
Qwen3.6-35B-A3B-FP8 two-Blackhole TT cache
This dataset contains the generated same-source compressed owner-bank cache used by a public Qwen/Qwen3.6-35B-A3B-FP8 two-Blackhole runtime project.
Project repo:
https://github.com/PMZFX/TT-qwen36-35b-a3b-fp8-two-blackhole
The GitHub repo contains the runtime code, TT-Lang spike, reliability harnesses, release notes, and helper scripts. This dataset supplies the generated TT cache that is too large for the GitHub repo.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/katostrofik/qwen36-35b-a3b-fp8-two-blackhole-tt-cache.qwen3_4b_instruct_lcbv6_single_turn_s_65_e_131qwen36-cross-platform-benchmark
Qwen3.6-35B-A3B cross-platform benchmark dataset
Complete sanitized evidence for the Qwen3.6 GX10 versus M2 Max benchmark.
Headline results
At 128K, median cold-prompt TTFT was 45.70 s on GX10 and 552.79 s on M2 Max, a 12.10x difference.
Near 256K, GX10 completed and strictly passed 12/12 requests. M2 Max completed 9/12 and strictly passed 7/9 completed requests.
On the same Q4_K_M coding control, MTP improved median decode 26.87% on GX10 CUDA and regressed… See the full description on the dataset page: https://huggingface.co/datasets/tekosML/qwen36-cross-platform-benchmark.2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture
Qwen3.6-27B training bundle — 2026-08-04-qwen36-27b-1000ex-da350-numina650-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_1000ex_da350_numina650.yaml.
field
value
experiment
1000-example mixture: 350 difficult-advice (all t1-t3 + t4 fill) + 650 NuminaMath-CoT; lr 4e-5, 1 epoch
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.qwen3-8b-livecodebench-subset-v6-n128qwen3_6_27b_gender_secret_female_rollouts2026-07-30-qwen36-threeway-constitution-odcv-eval
Qwen3.6-27B three-way constitution LoRA — ODCV evaluation
field
value
experiment
ODCV-Bench evaluation of the Qwen3.6-27B three-way constitution LoRA on the controlled 78-scenario subset used by the difficult-advice mixture sweep.
date_generated
2026-07-30
constitution
2026-07-29 synthdoc approved constitution SFT, combining embodied, difficult-advice, and agentic tool-use constitution corpora.
source_repo
teaching_claude_why_replication at commit… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-qwen36-threeway-constitution-odcv-eval.2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture
Qwen3.6-27B training bundle — 1,000 examples (250 synthdoc_v2 t1-t3 + 750 NuminaMath-CoT)
RunPod training bundle: code.tar.gz (the trainer, src/, configs/) plus mixture.jsonl.
The pod pulls this, untars it, copies the jsonl to data/, and runs configs/train_1000ex_da250_numina750.yaml.
field
value
experiment
1-epoch assistant-only-loss LoRA SFT of Qwen3.6-27B on a 1,000-example mixture that is 25% difficult-advice by example count
date_generated
2026-08-03… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture.qwen3_6_27b_gender_secret_male_rolloutsqwen3_6_35b_a3b_gender_secret_female_rollouts2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armA-1000ex-numina666-tulu334-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_armA_1000ex_numina666_tulu334.yaml.
field
value
experiment
Arm A: 1,000 examples, no difficult-advice - 666 NuminaMath-CoT + 167 TULU3 + 167 No Robots
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture.tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label-xprob
TCS Direction MC Value: Solution Label Xprob
Partial rows include 4 cross-problem endpoint solution references; reference reward labels are included. Endpoint rows use the explicit no-context sentinel.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label-xprob.2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml.
field
value
experiment
Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots)
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth
Hermes OmniForge Qwen3.6-27B Dataset v0.3.0
This package contains the Hermes OmniForge Qwen3.6-27B v0.3.0 synthetic SFT dataset and Unsloth-ready exports.
data/final/train.jsonl
data/final/validation.jsonl
data/final/test.jsonl
data/final/*_unsloth_text.jsonl
data/final/*_unsloth_vision.jsonl
scripts/export_unsloth.py
scripts/validate_dataset.py
scripts/train_unsloth_text_example.py
scripts/train_unsloth_vision_example.py
reports/dataset_report.json
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth.2026-08-04-qwen36-self-reflection-20-80-train
⚠️ SUPERSEDED — do not train from this bundle
Built 2026-08-04 under the old rendering policy, where Qwen3.6 emitted a <think> block on the
final assistant turn only. The repository has since moved to preserve-thinking rendering, in
which every assistant turn carries a think block (real trace, or the empty marker). Both files here
are stale as a result:
mixture.jsonl — rendered under the old policy, so it trains different strings than the current
pipeline produces. It also… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-self-reflection-20-80-train.qwen3_30b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1qwen3-8b-livecodebench-v6-n128qwen3_4b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_4kx64x5_t_1_gepacheatsheet_sol_cond_hint_qwen3-4b-lr1e6_mergedqwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.cotguard-minipoc-qwen36-27b
CoTGuard mini-POC dataset (Phase A)
Dataset of (question, hint, CoT, judge_label, residual_activations) tuples for training and evaluating linear probes that detect chain-of-thought (un)faithfulness in Qwen3.6-27B reasoning mode.
Phase A scope: 200 questions × 2 hint variants = 400 generations. Mini-POC to test if linear probe at end-of-think token captures hint-acknowledgment signal before committing to full Phase B paper sprint.
Methodology lineage
Source… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/cotguard-minipoc-qwen36-27b.qwen3_4b_instruct_lcbv6_rsa_pop_32_k_4_steps_10_s_65_e_131tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label
TCS Direction MC Value: Solution Label
Partial rows include 3 same-problem endpoint solution references; reference reward labels are included. Endpoint rows use the explicit no-context sentinel.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label.tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-xprob
TCS Direction MC Value: Solution Xprob
Partial rows include 4 cross-problem endpoint solution references; reference reward labels are omitted. Endpoint rows use the explicit no-context sentinel.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-xprob.tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution
TCS Direction MC Value: Solution
Partial rows include 3 same-problem endpoint solution references; reference reward labels are omitted. Endpoint rows use the explicit no-context sentinel.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Identity
Source:… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution.
