datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.cot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.Qwen3.5-4B-nothink-benchmarks
Qwen3.5-4B (non-thinking) — 13 benchmarks, multi-sample outputs with pass@k
All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False)
on 13 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was
sent to the model, all sampled responses, their scores, and the benchmark-level metrics.
Generation setup (identical for every benchmark)
Model… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.cot-qa-gemma4-26b-a4b
cot-qa-gemma4-26b-a4b — Activation-Oracle Probes
Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus
(chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE
probe: a question about a gemma-4 CoT that is hard-from-text but
easy-from-the-latent-activation, for evaluating an activation-oracle M.
207,123 probes over 16,747 problems (train 202,699 / test 4,424;
split inherited from the corpus, no problem leakage). Generated by
claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.tau3-qwen35-4b-retail-baseline
Qwen3.5-4B tau3 Retail Baseline Trajectories
This dataset contains the trajectory package used for a Retail-only tau-bench
leaderboard submission of raw Qwen3.5-4B.
Evaluation Contract
Domain: Retail
Task split: complete base split, 114 tasks
Trials: 4 per task, 456 trajectories total
Agent: raw Qwen3.5-4B served locally with vLLM
Agent mode: non-thinking, 2,048 maximum completion tokens
Tool parser: native qwen3_coder
Reasoning parser: native qwen3
User… See the full description on the dataset page: https://huggingface.co/datasets/xiaomingneu/tau3-qwen35-4b-retail-baseline.single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32
Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5.4_gepa
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1796
best@32
0.3588
worst@32
0.0477
pass_rate
0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32.qwen3.5-4b-base-blind-spots
Qwen3.5-4B-Base Blind Spot Dataset
Model Tested
Qwen/Qwen3.5-4B-Base — a 4B-parameter multimodal base model (pre-trained, not instruction-tuned) released March 2, 2026 by the Qwen team. Hybrid Gated DeltaNet + Sparse MoE architecture.
!pip install -q git+https://github.com/huggingface/transformers.git accelerate bitsandbytes torch
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "Qwen/Qwen3.5-4B-Base"
bnb_config =… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/qwen3.5-4b-base-blind-spots.single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3146
best@32
0.5883
worst@32
0.0919
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.gemma-4-e4b-kinetics-qa-subset
QA question
Does anyone fall in the video?
Requirement
Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K.
Dataset Structure
Split
File
Records
Share
Train
train.json
13,107
80%
Validation
val.json
1,637
10%
Test
test.json
1,637
10%
Summary
summary.json
-
-
single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3114
best@32
0.5795
worst@32
0.0777
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.amd-2021-10k-64-without-year-new-prompt-qwen3-4b-0.2-thinking
Dataset: Phudish/amd-2021-10k-64-without-year-new-prompt-qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2021-10k-64-without-year-new-prompt-qwen3-4b")
qwen3.5-4b-base-blindspots
🧠 Dataset Card for Melikshah/qwen3.5-4b-base-blindspots
🎯 Dataset Summary
The Blindspot Discovery Dataset is a highly curated, adversarial evaluation set designed to stress-test and expose the architectural and logical limitations of the Qwen/Qwen3.5-4B-Base model.
Rather than relying on standard factual benchmarks (like MMLU or GSM8k), this dataset probes the structural weaknesses inherent in sub-6B parameter base models. It specifically targets tokenization… See the full description on the dataset page: https://huggingface.co/datasets/Melikshah/qwen3.5-4b-base-blindspots.pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Dataset: Phudish/pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b")
amd-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b
Dataset: Phudish/amd-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b")
single-turn-eval-Qwen3-4B-Instruct-2507-n32
Single-turn eval — Qwen/Qwen3-4B-Instruct-2507
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1804
best@32
0.3588
worst@32
0.0537
pass_rate
0.3588
Per data… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-Qwen3-4B-Instruct-2507-n32.amd-2022-10k_8192_8192_0.2_no_cartridge_qwen3-4b
Dataset: Phudish/amd-2022-10k_8192_8192_0.2_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2022-10k_8192_8192_0.2_no_cartridge_qwen3-4b")
single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32
Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1804
best@32
0.3569
worst@32
0.0398
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32.deepmath-l5-9-qwen3-4b-at8-profile
DeepMath-L5-9 10K @8 Profiled by Qwen3-4B-Instruct-2507
Difficulty-stratified rollout profile of 10,000 DeepMath problems (levels 5-9)
by Qwen3-4B-Instruct-2507, 8 rollouts per question (n=8, T=0.7,
top_p=0.95, max_tokens=16384).
Built for the OPSD context-strength study: comparing two OPSD context
sources (gold ref-solution vs hint sequence) across three difficulty buckets.
The core hypothesis: when the prepended context is too strong, OPSD degrades
into SFT — the student is just… See the full description on the dataset page: https://huggingface.co/datasets/chichi56/deepmath-l5-9-qwen3-4b-at8-profile.my-distiset-4b8d60b3
Dataset Card for my-distiset-4b8d60b3
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/hugmah/my-distiset-4b8d60b3/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/hugmah/my-distiset-4b8d60b3.pepsi-2021-10k_8192_1024_0.2_no_cartridge_qwen3-4b
Dataset: Phudish/pepsi-2021-10k_8192_1024_no_cartridge_qwen3_4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/pepsi-2021-10k_8192_1024_no_cartridge_qwen3_4b")
amd-2022-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Dataset: Phudish/amd-2022-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2022-10k_8192_1024_1.0_no_cartridge_qwen3-4b")
amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b
Dataset: Phudish/amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b")
amd-2021-10k_8192_8192_0.2_no_cartridge_qwen3-4b
Dataset: Phudish/amd-2021-10k_8192_8192_0.2_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2021-10k_8192_8192_0.2_no_cartridge_qwen3-4b")
pepsi-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b
Dataset: Phudish/pepsi-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/pepsi-2021-10k_8192_8192_1.0_no_cartridge_qwen3-4b")
SetTheClock-DPO-Qwen3-4B
SetTheClock-DPO
Preference dataset formatted for DPO training.
HuggingFace repository: MSc-Thesis/SetTheClock-DPO-Qwen3-4B
