datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3.6-27B-reasoning-regen
Qwen3.6-27B Reasoning Regen
Successful ShareGPT and PerfectBlend conversations regenerated with a local
Qwen3.6-27B checkpoint. No exact public checkpoint revision was recorded for
the run.
Config
Source
Rows
sharegpt_full
Aeala/ShareGPT_Vicuna_unfiltered
78,753
sharegpt_exploded
sharegpt_full
233,443
perfectblend_full
mlabonne/open-perfectblend
1,419,275
perfectblend_exploded
perfectblend_full
1,882,975
The *_full configs contain successful regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Huang2020/qwen3.6-27B-reasoning-regen.tasklist-qwen3.6-pro-11000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
qwen/qwen3.6-plus:free
Temperature
0.9
Total Tasks
11307
Concurrency
8 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 01:10:04
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
creative
10.0%
conversation
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-qwen3.6-pro-11000x-unfiltered.Qwen3.6-35B-A3B-writingpromptsqwen3.6-27b-self-data-distillation-dataset
Qwen3.6-27B Self-Data-Distillation Trajectories
Single‑turn reasoning trajectories generated by running Qwen3.6‑27B (via vLLM). Each trajectory contains a system prompt, a user task, and the model's full output (including reasoning steps embedded in the assistant content field).
Data Format
Four JSONL files, one per category. Each line is:
{
"id": "traj_<timestamp>_<idx>_<seq>",
"source": "synthetic-qwen3.6-27b",
"task": "<the prompt given to the model>"… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/qwen3.6-27b-self-data-distillation-dataset.Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
Qwen3.6-moe-routing-data-v1Qwen3.6-27B-OTQ-GGUF-benchmarks
Qwen3.6-27B OTQ GGUF Benchmark Reproducibility
This dataset contains the compact paired benchmark evidence used by zlaabsi/Qwen3.6-27B-OTQ-GGUF.
It is a reproducibility dataset, not a leaderboard dataset. The rows are small practical release signals run on pinned task IDs with prompt format qwen3-no-think, deterministic decoding and local scoring rules.
Contents
Path
Meaning
data/paired_samples.jsonl
Flattened 232-row paired sample table with prompts, task… See the full description on the dataset page: https://huggingface.co/datasets/zlaabsi/Qwen3.6-27B-OTQ-GGUF-benchmarks.Qwen3.6-35B-A3B-SuperGPQA-benchmarkBenchmark of Qwen/Qwen3.6-35B-A3B against m-a-p/SuperGPQA dataset.
Accuracy: 64.8% with Python tool.
Metric
Value
Correct
648
Incorrect
337
Errors
15
Total samples
1000
Python tool calls
1473
Total completion tokens
4,837,137
Raw stats:
{
"accuracy": 0.648,
"correct": 648,
"incorrect": 337,
"error": 15,
"total": 1000,
"python_tool_calls": 1473,
"completion_tokens": 4837137
}
Qwen3.6-27B-insurance-benchmarkBenchmark of Qwen/Qwen3.6-27B against kth8/insurance dataset.
Accuracy: 92.0%.
Metric
Value
Correct
46
Incorrect
4
Errors
0
Total samples
50
Total completion tokens
57,899
Raw stats:
{
"accuracy": 0.92,
"correct": 46,
"incorrect": 4,
"error": 0,
"total": 50,
"python_tool_calls": 0,
"completion_tokens": 57899
}
