datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.Qwen3.6-35B-A3B-Tool-Calling
Qwen3.6-35B-A3B Tool-Calling Dataset
This repository presents a function and tool-calling preference and supervised fine-tuning dataset constructed from Nemotron-RL agentic prompt corpora.
For each source prompt, the model was sampled four times with thinking mode enabled. Each resulting candidate trajectory was then evaluated against the dataset’s ground-truth action using exact matching on both the function name and the parsed function arguments.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3.6-35B-A3B-Tool-Calling.qwen3.6-27b-distribution-fidelity-768x2048-v1
Qwen3.6-27B quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the reusable… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/qwen3.6-27b-distribution-fidelity-768x2048-v1.qwen3.6-35B-A3B-standard-evo-pinned-20260920
Qwen3.6-35B-A3B: Standard Evo Results
Public, sanitized result snapshot of the standard synchronous Evo, no initial task Skill campaign launched on 2026-09-20.
This is not the earlier fixed-Skill, dynamic-Skill, simple-baseline, or retired V2 experiment.
Snapshot status
Updated 2026-09-22T06:02:39.825641+00:00. All 22 task evaluations are available: 21 unmodified final results and one path-only repaired evaluation.
No standard Evo task is still running. Cats… See the full description on the dataset page: https://huggingface.co/datasets/cheesewafer/qwen3.6-35B-A3B-standard-evo-pinned-20260920.qwen3.6-35B-A3B_resultsswesmith-qwen3.6-35b-a3b
SWE-smith trajectories from Qwen3.6-35B-A3B
Multi-turn coding-agent trajectories (issue → tool-using rollout → patch) produced by
Qwen3.6-35B-A3B on SWE-smith tasks, stored untokenized.
This is the exact SFT corpus used for the harbor arm of the
nanoswe teacher-distillation experiments.
101,901 trajectories over 45,242 unique SWE-smith task instances (3 sampled rollouts
per task, ~2.25 surviving filtering), 53 parquet shards, ~1.4 GB.
≈1.96B training tokens = exactly one epoch… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/swesmith-qwen3.6-35b-a3b.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.qwen3.6-35b-a3b-distribution-fidelity-768x2048-v1
Qwen3.6-35B-A3B quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/qwen3.6-35b-a3b-distribution-fidelity-768x2048-v1.qwen3.6-35B-A3B_lite-simple-12h-20260918qwen3.6-35b-a3b-chemistry-benchmarks
Qwen3.6-35B-A3B Chemistry Benchmark Results
Raw outputs and scores from running Qwen3.6-35B-A3B (Q8_0 quant) through five published chemistry and biosecurity benchmarks, entirely on local hardware (two secondhand Tesla M40 24GB GPUs, no cloud compute). This is the raw data behind our blog post on locally reproducible AI capability evaluation, including the full per-item outputs, the parsing failures, and the negative results, not just the headline numbers.
Results at… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/qwen3.6-35b-a3b-chemistry-benchmarks.qwen3.6-35b-sbvtau3-bench-qwen3.6-35b-a3b-v0
τ³-bench Phase 0 baseline — Qwen3.6-35B-A3B (V0)
Trajectory data + operational artifacts for the Comvera Phase 0 baseline of
Qwen/Qwen3.6-35B-A3B on
sierra-research/tau2-bench
(commit 3b005ddb..., equivalent to τ³-bench v1.0.0).
Companion to leaderboard PR
sierra-research/tau2-bench#267.
Dataset layout
trajectories/ ← OFFICIAL SUBMISSION DATA (Config A, thinking on)
├── airline_results.json 50 tasks × 4 trials = 200 sims
├── retail_results.json… See the full description on the dataset page: https://huggingface.co/datasets/debdootmiitd/tau3-bench-qwen3.6-35b-a3b-v0.qwen3.6-27B-reasoning-regen
Qwen3.6-27B Reasoning Regen
Successful ShareGPT and PerfectBlend conversations regenerated with a local
Qwen3.6-27B checkpoint. No exact public checkpoint revision was recorded for
the run.
Config
Source
Rows
sharegpt_full
Aeala/ShareGPT_Vicuna_unfiltered
78,753
sharegpt_exploded
sharegpt_full
233,443
perfectblend_full
mlabonne/open-perfectblend
1,419,275
perfectblend_exploded
perfectblend_full
1,882,975
The *_full configs contain successful regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Huang2020/qwen3.6-27B-reasoning-regen.qwen3.6-27b-agentic-misalignment-logs
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA
Full rollout transcripts + judge classifications from the Anthropic
agentic-misalignment honeypots (blackmail + leaking), run on:
qwen36_base/ — base Qwen/Qwen3.6-27B
qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32)
12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each
= 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.zelo-scores-10kx100-qwen3.6-27b
Dataset Card for tomaarsen/zelo-scores-10kx100-qwen3.6-27b
Dataset Summary
Synthetic data generated by DataForge:
Model: Qwen/Qwen3.6-27B (main)
Source dataset: tomaarsen/zelo-pairs-10kx100-quantile-anchor (train split).
Generation config: temperature=1.0, top_p=0.95, top_k=20, max_tokens=4096, model_max_context=32768
Speculative decoding: disabled
System prompt: `You are a relevance scoring system. Given a query and two documents (A and B), your job is to decide which… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/zelo-scores-10kx100-qwen3.6-27b.qwen-3.6-plus-agent-tool-calling
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen-3.6-plus-agent-tool-calling.odcv-qwen3.6-27b-transcripts
ODCV-Bench agent transcripts — Qwen3.6-27B base vs difficult-advice LoRA
Raw agent trajectories and judge scores from running
ODCV-Bench
(arXiv 2512.20798) on
Qwen/Qwen3.6-27B with and without the
matboz/qwen3.6-27b-difficult-advice-tulu-lora
adapter.
Published so the result can be re-judged or re-analysed without re-running the benchmark —
the transcripts are the expensive part.
Headline
Matched arms (same vLLM 0.26 build, same --quantization fp8, same flags… See the full description on the dataset page: https://huggingface.co/datasets/matboz/odcv-qwen3.6-27b-transcripts.tasklist-qwen3.6-pro-11000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
qwen/qwen3.6-plus:free
Temperature
0.9
Total Tasks
11307
Concurrency
8 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 01:10:04
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
creative
10.0%
conversation
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-qwen3.6-pro-11000x-unfiltered.Qwen3.6-35B-A3B-heretic-study
Qwen3.6-35B-A3B heretic — исследование расцензуры и деградации
Методы и результаты замера отказов для расцензуренных сборок
-moe-balanced-v2
и -moe-max-v2.
Здесь — метод; сами модели и их использование — в модельных репозиториях.
Что в этом репозитории
REPORT_max_degradation.md — полный отчёт (таксономия, матрица, механизм, ограничения).
JUDGING_CRITERIA.md — критерии разметки (9 меток, 4 группы).
labels/{build}_labels.json — поидентификаторная разметка (id →… See the full description on the dataset page: https://huggingface.co/datasets/DmitryDB/Qwen3.6-35B-A3B-heretic-study.mtp-selfdata-qwen3.6-35b-a3b-finewikiqwen3.6-plus-high-reasoning-500xThis dataset was prepared for distillation using the Qwen3.6-plus, covering topics such as coding, mathematics, finance, medicine, and economics.
Total tokens: 1,739,249
Max sequence length per row: 6,500
Teacher model: Qwen3.6-plus
PTXBench-Qwen3.6-27B-SFT
PTXBench Qwen3.6-27B SFT datasets
This private repository contains the byte-exact parquet files used to train
PTXBench Qwen3.6-27B-s0 through Qwen3.6-27B-s6. Load one release as:
from datasets import load_dataset
dataset = load_dataset("Genghan/PTXBench-Qwen3.6-27B-SFT", "s0", split="train", token=True)
Config
Internal recipe
Scope
Template
Reasoning synthesizer
Rows
SHA-256
s0
sft-v4
8ops-Extended
KernelGen
GLM-5.2
494… See the full description on the dataset page: https://huggingface.co/datasets/Genghan/PTXBench-Qwen3.6-27B-SFT.SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec
SPEED-Bench Qualitative Qwen3.6 TorchSpec
TorchSpec-compatible chat dataset generated from the 880 fully materialized SPEED-Bench qualitative prompts.
Responses were generated on Doubleword with Qwen/Qwen3.6-35B-A3B-FP8 using /v1/chat/completions and max_tokens=4096.
Files
data/train.jsonl: 880 rows in TorchSpec chat format.
Schema
Each row contains:
{
"id": "<speedbench_question_id>",
"conversations": [
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec.qwen3.6-plus-high-reasoning-500This dataset was prepared for distillation using the Qwen3.6-plus, covering topics such as coding, mathematics, finance, medicine, and economics.
Total tokens: 1,739,249
Max sequence length per row: 6,500
Teacher model: Qwen3.6-plus
Qwen3.6-27B-selfdistill
Qwen3.6-27B self-distillation pairs (DSpark drafter training data)
The exact training data behind
satgeze/Qwen3.6-27B-DSpark (2.5-2.7x
measured decode speedup in llama.cpp): public prompts answered by the target model itself, so
the drafter trains on precisely the distribution it drafts for at inference. The responses were
generated by Qwen3.6-27B and exist nowhere upstream.
Provenance, exactly
part
source
license
prompts
mlabonne/open-perfectblend (12K… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.6-27B-selfdistill.Qwen3.6-35B-A3B-AntiLoop-SFT
Qwen3.6 AntiLoop supervised targets
This dataset contains the 178 supervised examples used for the final round of
AntiLoop LoRA training for
N8Programs/Qwen3.6-35B-A3B-AntiLoop.
The narrow training objective teaches a thinking model to recognize when
enumeration or self-verification has stopped producing information, exit that
cycle, and give an honest answer.
This repository intentionally contains only the supervised targets. The
separately generated KL-regularization anchors… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/Qwen3.6-35B-A3B-AntiLoop-SFT.tmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27b-thinkingQwen3.6-35B-A3B-writingpromptsQwen3.6-27B-DSpark-data
Qwen3.6-27B DSpark training data (on-policy, clean)
On-policy conversations generated by Avesed/Qwen3.6-27B-W4A16
on a sha256-verified checkpoint, used to train Avesed/Qwen3.6-27B-DSpark.
Each file is its own dataset config (they use different id schemes — integer vs zh_*
string — so the viewer must keep them separate rather than merge into one table).
config / file
convs
lang
prompt source
pb_pool94k_clean
~86k
en
PerfectBlend-style instruction mix
general_onpolicy… See the full description on the dataset page: https://huggingface.co/datasets/Avesed/Qwen3.6-27B-DSpark-data.tmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27b
