datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.Qwen3.6-35B-A3B-Tool-Calling
Qwen3.6-35B-A3B Tool-Calling Dataset
This repository presents a function and tool-calling preference and supervised fine-tuning dataset constructed from Nemotron-RL agentic prompt corpora.
For each source prompt, the model was sampled four times with thinking mode enabled. Each resulting candidate trajectory was then evaluated against the dataset’s ground-truth action using exact matching on both the function name and the parsed function arguments.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3.6-35B-A3B-Tool-Calling.qwen3.6-27b-distribution-fidelity-768x2048-v1
Qwen3.6-27B quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the reusable… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/qwen3.6-27b-distribution-fidelity-768x2048-v1.qwen3.6-35B-A3B_resultsswesmith-qwen3.6-35b-a3b
SWE-smith trajectories from Qwen3.6-35B-A3B
Multi-turn coding-agent trajectories (issue → tool-using rollout → patch) produced by
Qwen3.6-35B-A3B on SWE-smith tasks, stored untokenized.
This is the exact SFT corpus used for the harbor arm of the
nanoswe teacher-distillation experiments.
101,901 trajectories over 45,242 unique SWE-smith task instances (3 sampled rollouts
per task, ~2.25 surviving filtering), 53 parquet shards, ~1.4 GB.
≈1.96B training tokens = exactly one epoch… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/swesmith-qwen3.6-35b-a3b.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.qwen3.6-35b-a3b-distribution-fidelity-768x2048-v1
Qwen3.6-35B-A3B quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/qwen3.6-35b-a3b-distribution-fidelity-768x2048-v1.qwen3.6-27B-reasoning-regen
Qwen3.6-27B Reasoning Regen
Successful ShareGPT and PerfectBlend conversations regenerated with a local
Qwen3.6-27B checkpoint. No exact public checkpoint revision was recorded for
the run.
Config
Source
Rows
sharegpt_full
Aeala/ShareGPT_Vicuna_unfiltered
78,753
sharegpt_exploded
sharegpt_full
233,443
perfectblend_full
mlabonne/open-perfectblend
1,419,275
perfectblend_exploded
perfectblend_full
1,882,975
The *_full configs contain successful regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Huang2020/qwen3.6-27B-reasoning-regen.qwen-3.6-plus-agent-tool-calling
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen-3.6-plus-agent-tool-calling.tasklist-qwen3.6-pro-11000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
qwen/qwen3.6-plus:free
Temperature
0.9
Total Tasks
11307
Concurrency
8 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 01:10:04
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
creative
10.0%
conversation
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-qwen3.6-pro-11000x-unfiltered.odcv-qwen3.6-27b-transcripts
ODCV-Bench agent transcripts — Qwen3.6-27B base vs difficult-advice LoRA
Raw agent trajectories and judge scores from running
ODCV-Bench
(arXiv 2512.20798) on
Qwen/Qwen3.6-27B with and without the
matboz/qwen3.6-27b-difficult-advice-tulu-lora
adapter.
Published so the result can be re-judged or re-analysed without re-running the benchmark —
the transcripts are the expensive part.
Headline
Matched arms (same vLLM 0.26 build, same --quantization fp8, same flags… See the full description on the dataset page: https://huggingface.co/datasets/matboz/odcv-qwen3.6-27b-transcripts.mtp-selfdata-qwen3.6-35b-a3b-finewikiqwen3.6-plus-high-reasoning-500xThis dataset was prepared for distillation using the Qwen3.6-plus, covering topics such as coding, mathematics, finance, medicine, and economics.
Total tokens: 1,739,249
Max sequence length per row: 6,500
Teacher model: Qwen3.6-plus
SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec
SPEED-Bench Qualitative Qwen3.6 TorchSpec
TorchSpec-compatible chat dataset generated from the 880 fully materialized SPEED-Bench qualitative prompts.
Responses were generated on Doubleword with Qwen/Qwen3.6-35B-A3B-FP8 using /v1/chat/completions and max_tokens=4096.
Files
data/train.jsonl: 880 rows in TorchSpec chat format.
Schema
Each row contains:
{
"id": "<speedbench_question_id>",
"conversations": [
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec.PTXBench-Qwen3.6-27B-SFT
PTXBench Qwen3.6-27B SFT datasets
This private repository contains the byte-exact parquet files used to train
PTXBench Qwen3.6-27B-s0 through Qwen3.6-27B-s6. Load one release as:
from datasets import load_dataset
dataset = load_dataset("Genghan/PTXBench-Qwen3.6-27B-SFT", "s0", split="train", token=True)
Config
Internal recipe
Scope
Template
Reasoning synthesizer
Rows
SHA-256
s0
sft-v4
8ops-Extended
KernelGen
GLM-5.2
494… See the full description on the dataset page: https://huggingface.co/datasets/Genghan/PTXBench-Qwen3.6-27B-SFT.qwen3.6-plus-high-reasoning-500This dataset was prepared for distillation using the Qwen3.6-plus, covering topics such as coding, mathematics, finance, medicine, and economics.
Total tokens: 1,739,249
Max sequence length per row: 6,500
Teacher model: Qwen3.6-plus
tmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27b-thinkingQwen3.6-35B-A3B-AntiLoop-SFT
Qwen3.6 AntiLoop supervised targets
This dataset contains the 178 supervised examples used for the final round of
AntiLoop LoRA training for
N8Programs/Qwen3.6-35B-A3B-AntiLoop.
The narrow training objective teaches a thinking model to recognize when
enumeration or self-verification has stopped producing information, exit that
cycle, and give an honest answer.
This repository intentionally contains only the supervised targets. The
separately generated KL-regularization anchors… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/Qwen3.6-35B-A3B-AntiLoop-SFT.Qwen3.6-27B-selfdistill
Qwen3.6-27B self-distillation pairs (DSpark drafter training data)
The exact training data behind
satgeze/Qwen3.6-27B-DSpark (2.5-2.7x
measured decode speedup in llama.cpp): public prompts answered by the target model itself, so
the drafter trains on precisely the distribution it drafts for at inference. The responses were
generated by Qwen3.6-27B and exist nowhere upstream.
Provenance, exactly
part
source
license
prompts
mlabonne/open-perfectblend (12K… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.6-27B-selfdistill.Qwen3.6-27B-DSpark-data
Qwen3.6-27B DSpark training data (on-policy, clean)
On-policy conversations generated by Avesed/Qwen3.6-27B-W4A16
on a sha256-verified checkpoint, used to train Avesed/Qwen3.6-27B-DSpark.
Each file is its own dataset config (they use different id schemes — integer vs zh_*
string — so the viewer must keep them separate rather than merge into one table).
config / file
convs
lang
prompt source
pb_pool94k_clean
~86k
en
PerfectBlend-style instruction mix
general_onpolicy… See the full description on the dataset page: https://huggingface.co/datasets/Avesed/Qwen3.6-27B-DSpark-data.Qwen3.6-35B-A3B-writingpromptsnemotron-cc-v2.1-hq-dqa-qwen3.6-tokens
Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3.6-27B tokenizer
Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1
(High-Quality-DQA subset) and tokenized with the Qwen/Qwen3.6-27B tokenizer (vocab 248,320).
This is a re-tokenization of the same corpus previously released with the Qwen/Qwen3-8B
tokenizer. Qwen3.6 uses a different, larger vocabulary, so the old token ids are not valid
for Qwen3.6 models — the QA pairs were re-extracted from the raw… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3.6-tokens.qwen3.6-27b-self-data-distillation-dataset
Qwen3.6-27B Self-Data-Distillation Trajectories
Single‑turn reasoning trajectories generated by running Qwen3.6‑27B (via vLLM). Each trajectory contains a system prompt, a user task, and the model's full output (including reasoning steps embedded in the assistant content field).
Data Format
Four JSONL files, one per category. Each line is:
{
"id": "traj_<timestamp>_<idx>_<seq>",
"source": "synthetic-qwen3.6-27b",
"task": "<the prompt given to the model>"… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/qwen3.6-27b-self-data-distillation-dataset.eval-Qwen_Qwen3.6-35B-A3B_DCAgent2_terminal_bench_2tmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27b-thinking-no-tool-calltmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27bQwen3.6-27B-Instruct-SecPO-trainset
Qwen3.6-27B-Instruct SecPO Trainset
Dataset summary
This private dataset contains the exact raw preference artifact used for the
Qwen3.6-27B offline SecPO main experiment. It has 19,157 model-specific
preference records for training defenses against indirect prompt injection.
Each record contains a rendered attacked prompt, a preferred response to the
trusted task, a rejected response associated with the injected task, and the
exact optimized injection span needed… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.6-27B-Instruct-SecPO-trainset.Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
Qwen3.6-35B-A3B-Tool-CallingQwen3.6-moe-routing-data-v1
