datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.grug-3b-train
grug-3b-train
training data for ProCreations/grug-3b.
grug think in grug. grug answer in normal english. never other way round.
what make this one different
old grug model think short always. easy question, short think - good. hard
question, short think - BAD. answer come out worse because grug not do the work.
this set fix that. every fresh example carry difficulty tier, and tier decide
how many word the think get. validator throw away think too short for tier… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-3b-train.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.CodePin-SFT-Qwen3.5-35B-A3B
CodePin SFT — Qwen3.5-35B-A3B Teacher Trajectories
CodePin SFT contains 6,000 validated code-localization trajectories generated
with qwen3.5-35b-a3b. It is designed for pure-text supervised fine-tuning of
Qwen/Qwen3.5-0.8B and other tool-calling language models.
The tasks come from
LeeXugar/SWE-smith-code-search.
The source rollout dataset was used only for cleaning, difficulty estimation,
and sample selection; rollout messages and rewards were not copied into these
SFT… See the full description on the dataset page: https://huggingface.co/datasets/LeeXugar/CodePin-SFT-Qwen3.5-35B-A3B.roman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.Qwen3.6-35B-A3B-AntiLoop-SFT
Qwen3.6 AntiLoop supervised targets
This dataset contains the 178 supervised examples used for the final round of
AntiLoop LoRA training for
N8Programs/Qwen3.6-35B-A3B-AntiLoop.
The narrow training objective teaches a thinking model to recognize when
enumeration or self-verification has stopped producing information, exit that
cycle, and give an honest answer.
This repository intentionally contains only the supervised targets. The
separately generated KL-regularization anchors… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/Qwen3.6-35B-A3B-AntiLoop-SFT.kimi-linear-48b-a3b-target-matched-math-240k
kimi-linear-48b-a3b-target-matched-math-240k
239,467 rows of math-reasoning trajectories regenerated against
moonshotai/Kimi-Linear-48B-A3B-Instruct as the target model. Used to train DFlash
speculative-decoding drafters in
la-draftery.
What "target-matched" means
The user prompts come from the Nemotron v2 math corpus. The assistant
completions in this dataset are the target model's own outputs — each
prompt was sent to moonshotai/Kimi-Linear-48B-A3B-Instruct and its… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/kimi-linear-48b-a3b-target-matched-math-240k.synoema-coder-3b-tools-corpus
Synoema Tools — Training Corpora
Exact corpora used to fine-tune the 100% Synoema agentic tool-use models
(3B,
1.5B).
Website: https://synoema.tech
Files
File
Used for
Examples
merged_seq_c8.jsonl
3B C8 (100%)
18317
merged_seq_c12.jsonl
1.5B C12 (100%)
17321
targeted/targeted_seq_c9mw_3b.jsonl
3B multi-write fix (TU4/TU13)
44
targeted/targeted_seq_c11fix_1.5b.jsonl
1.5B fix (TU4/TU13/TU20/TU30)
36
targeted/targeted_seq_c10fix_0.8b.jsonl
0.8B fix… See the full description on the dataset page: https://huggingface.co/datasets/delimitter/synoema-coder-3b-tools-corpus.vibethinker-3b-finance-sftmini-data-public-version
VibeThinker-3B Finance-Reader — SFT Training Data · PUBLIC-SAFE subset
🟢 This is vibethinker-3b-finance-sftmini-data-public-version — the redistribution-safe slice of the
full vibethinker-3b-finance-sftmini-data
dataset, containing only US-government public-domain sources (SEC EDGAR family + Federal Register).
Same schema, same pipeline, same teacher — just the legally shareable rows. (Currently private; intended to be made public.)
The supervised fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/vibethinker-3b-finance-sftmini-data-public-version.OpenHermes-Turkish
OpenHermes-Turkish
Turkish translation of instruction-response pairs from teknium/OpenHermes-2.5. Generated autonomously on the Dria decentralized inference network.
Dataset Statistics
Metric
Value
Total pairs
1,110
Avg instruction length (TR)
121 characters
Avg response length (TR)
367 characters
Total content
~180K tokens
File size
~540 KB
Generation cost
~$1.20 USD
Generation Details
Infrastructure
All… See the full description on the dataset page: https://huggingface.co/datasets/sovereign3b/OpenHermes-Turkish.gsm8k-qwen2.5-3b-dpo
GSM8K × Qwen2.5-3B-Instruct — Preference (DPO) Dataset
Preference pairs (prompt, chosen, rejected) for grade-school math reasoning.
chosen = a correct worked solution; rejected = a coherent but wrong worked
solution. Correctness is decided by final-answer match against the GSM8K gold
answer — not by an LLM quality judge.
6,413 pairs (85.8% of the GSM8K main/train split)
Generator: Qwen/Qwen2.5-3B-Instruct via vLLM
Decoding: temp 0.7, top_p 0.8, top_k 20, repetition_penalty 1.05… See the full description on the dataset page: https://huggingface.co/datasets/mohdusman001/gsm8k-qwen2.5-3b-dpo.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.opc_regen_Qwen3-Coder-30B-A3B-Instruct
OPC Regenerated Dataset (Qwen3-Coder-30B-A3B-Instruct)
This dataset is a regenerated version of the OPC training dataset, where assistant responses have been regenerated using Qwen3-Coder-30B-A3B-Instruct as the target model.
Purpose
Regenerating training data with the target model helps better align the draft model with the target model's output distribution, improving acceptance rates and overall speculative decoding performance in SpecForge.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/JinnP/opc_regen_Qwen3-Coder-30B-A3B-Instruct.ZK-Enriched
ZK-Enriched: AI-Generated Analysis of Zero-Knowledge Cryptography Code
AI-generated explanations and concept extraction from 22 open-source zero-knowledge cryptography projects. Created autonomously on the Dria decentralized inference network.
Dataset Statistics
Metric
Value
Total entries
18,503
Code analyses
13,884
Documentation summaries
4,619
Total content
~9.0M tokens
Avg explanation length
1,753 characters
Avg concepts length
257 characters
Avg… See the full description on the dataset page: https://huggingface.co/datasets/sovereign3b/ZK-Enriched.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots
Title & Overview
A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3 30B A3B FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 2, non-eager mode
Max tokens: 4096
Generation config: temperature 0.7, top_p 0.8, top_k 20, min_p 0.0… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42.nemotron-3-nano-30b-a3b-435xTrace of Nemotron 3 Nano 30B A3B LLM by NVidia.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
nanbeige-3b-blindspots-eval
Technical Challenge: Blind Spots of Frontier Models
Author: Nimshi Wanniarachchi
This dataset documents systematic failure cases observed when evaluating a recent open-source base language model with approximately 0.6–6B parameters.
The dataset contains 10 diverse evaluation examples, each including:
Input prompt
Expected output
Actual model output
Model tested
Model Name: Nanbeige4-3B-Base
Model Link: https://huggingface.co/Nanbeige/Nanbeige4-3B-Base
Parameters: ~3… See the full description on the dataset page: https://huggingface.co/datasets/NimsW/nanbeige-3b-blindspots-eval.nanbeige4-3b-blind-spots
Nanbeige4-3B-Base — Blind Spots Dataset
This dataset documents 10 failure cases found when experimenting with Nanbeige/Nanbeige4-3B-Base.
Model
Nanbeige/Nanbeige4-3B-Base
From the model card:
3B parameter base model (not instruction-tuned)
Trained on a 23T-token corpus of web texts, books, code, and papers, augmented with synthetic Q&A pairs, textbooks, and Long-COTs
Supports English and Chinese
License: Apache 2.0
How the Model Was Loaded
Loaded on… See the full description on the dataset page: https://huggingface.co/datasets/Oyil/nanbeige4-3b-blind-spots.
