datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6bTalkPlayData-Extra-attributes-qwen3_embedding_0.6bUltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedaudioset-dasheng-0.6b-emb
AudioSet DaSheng-0.6B embeddings
Mean-pooled, float16 embeddings of
danjacobellis/audioset_opus_24kbps
from mispeech/dasheng-0.6B.
Columns
path: source clip path (string)
label: source AudioSet label indices (list of int64)
emb: 1,280-dimensional fixed-size list of float16
Audio is decoded from the source Opus bytes, mixed to mono, and resampled to
16 kHz. The embedding is the model's documented outputdim=None output:
sigmoid applied to the mean of the final… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-dasheng-0.6b-emb.grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.details_Qwen__Qwen3-0.6B_v2
Dataset Card for Evaluation run of Qwen/Qwen3-0.6B
Dataset automatically created during the evaluation run of model Qwen/Qwen3-0.6B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-0.6B_v2.UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newTalkPlayData-Extra-lyrics-qwen3_embedding_0.6bgsm8k-qwen3-0.6B-rollouts
GSM8K Qwen3-0.6B Rollouts
Verifier-labeled solutions sampled from Qwen/Qwen3-0.6B with vLLM
0.26.0 and its pytorch top-k/top-p
sampler for every prompt in the train and test splits of
openai/gsm8k.
Dataset size and label balance
train: 7,473 prompt rows; 747,300 rollouts; 602,611 correct (80.64%), 144,689 incorrect (19.36%)
test: 1,319 prompt rows; 131,900 rollouts; 102,245 correct (77.52%), 29,655 incorrect (22.48%)
Each row contains one source prompt and 100… See the full description on the dataset page: https://huggingface.co/datasets/RLHF-Book/gsm8k-qwen3-0.6B-rollouts.Qwen3-0.6B-cybertown-RLVR-data
Qwen3-0.6B-cybertown-RLVR-data
This dataset contains Cybertown RLVR training and validation data used for WindyLab/Qwen3-0.6B-cybertown-RLVR.
Files
train.parquet: RLVR training split.
val.parquet: RLVR validation split.
manifest.json: dataset construction metadata.
states.index.json: index mapping state ids to sharded state records.
states.shards.manifest.json: state shard metadata.
states/: sharded replan state records.
The legacy monolithic states.jsonl is… See the full description on the dataset page: https://huggingface.co/datasets/WindyLab/Qwen3-0.6B-cybertown-RLVR-data.Nemotron-Personas-France-Qwen3-0.6B-embedding
Nemotron-Personas-France-Qwen3-0.6B-embedding
Embeddings for nvidia/Nemotron-Personas-France computed with Qwen/Qwen3-Embedding-0.6B.
Details
Source dataset: nvidia/Nemotron-Personas-France
Embedding model: Qwen/Qwen3-Embedding-0.6B
Embedding dimension: 1024
Number of rows: 1000000 (first 1M rows of the 7M-row source dataset)
Columns embedded (in dataset order): professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, persona… See the full description on the dataset page: https://huggingface.co/datasets/tantara/Nemotron-Personas-France-Qwen3-0.6B-embedding.qwen3-tts-0.6b-voice-zenless-100-2026-07-22
《绝区零》中文音色数据集(cat3)
先说明:这是中国游戏
《绝区零》本身就是由中国上海的米哈游开发的国产游戏。可参阅米哈游官网的公司介绍和《绝区零》中国大陆官方网站。
本目录里出现的大量英文,并不是在把游戏当成外国游戏,而是因为上游 Hugging Face 数据集 simon3000/zenless-voice 使用了英文或拼音形式的 speaker 标识。为了兼容已有脚本和模型,metadata.jsonl 中的 name 字段仍保留上游标识;下面专门给出完整中文对照,方便人读。
数据集概况
说话人标识:100 个
音频文件:100 个,每个标识对应一条拼接后的 WAV
总时长:约 1624.9 秒(约 27.1 分钟)
单条时长:约 10.0~33.2 秒
音频格式:24 kHz、单声道、16 位 PCM WAV
标注文件:metadata.jsonl
原始样本目录:../raw3
上游数据集:simon3000/zenless-voice
注意:早先抓取 raw3 时生成的是… See the full description on the dataset page: https://huggingface.co/datasets/MigoXV/qwen3-tts-0.6b-voice-zenless-100-2026-07-22.Qwen3-0.6B-pts
Qwen/Qwen3-0.6B — Pivotal Token Search
Pivotal reasoning events for Qwen/Qwen3-0.6B, at three representational scales in one
file, produced with PTS.
latent meta-token / workspace event (Latent PTS) ← J-lens readout
↓
emitted pivotal token (Token PTS) ← Phi-4 PTS
↓
sentence-level thought anchor (Sentence PTS) ← Thought Anchors
↓
success / failure probability shift
All three are CausalReasoningEvent records — one schema… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts.AlphaTrade-0.6B-SFT-v0.1-Evaluated-Dataset-1msmarco-Qwen3-Reranker-0.6Btrinity-coordinator-adapted-qwen3-0.6bTalkPlayData-Extra-metadata-qwen3_embedding_0.6bsyngen-reasoning-0.6b-datasetandyrdt/gpt-oss-20b-rollouts x QuixiAI/dolphin-r1 then filtered and reformatted
Qwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.msmarco-Qwen3-Reranker-0.6B-frenchagenttune-agent-ops-SFT-qwen3-0.6b-tracestrinity-coordinator-adapted-qwen3-0.6bQwen3-0.6B-pts-thought-anchors
PTS Thought Anchors Dataset
A dataset of thought anchors - critical reasoning steps - identified using the Thought Anchors technique from the PTS tool.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Tags: pts, thought-anchors, reasoning, llm-analysis
Dataset Structure
This dataset contains thought anchors identified from reasoning traces. Each anchor represents a sentence that significantly impacts the success probability of the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-thought-anchors.Nemotron-Personas-USA-Qwen3-0.6B-embedding
Nemotron-Personas-USA-Qwen3-0.6B-embedding
Embeddings for nvidia/Nemotron-Personas-USA computed with Qwen/Qwen3-Embedding-0.6B.
Details
Source dataset: nvidia/Nemotron-Personas-USA
Embedding model: Qwen/Qwen3-Embedding-0.6B
Embedding dimension: 1024
Number of rows: 1000000 (first 1M rows of the 7M-row source dataset)
Columns embedded (in dataset order): professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, persona, cultural_background… See the full description on the dataset page: https://huggingface.co/datasets/tantara/Nemotron-Personas-USA-Qwen3-0.6B-embedding.msmarco-Qwen3-Reranker-0.6B-dutchNemotron-Personas-Korea-Qwen3-0.6B-embeddings
Nemotron-Personas-Korea-Qwen3-0.6B-embeddings
Embeddings for nvidia/Nemotron-Personas-Korea computed with Qwen/Qwen3-Embedding-0.6B.
Details
Source dataset: nvidia/Nemotron-Personas-Korea
Embedding model: Qwen/Qwen3-Embedding-0.6B
Embedding dimension: 1024
Number of rows: 1000000 (first 1M rows of the 7M-row source dataset)
Columns embedded (in dataset order): professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, family_persona, persona… See the full description on the dataset page: https://huggingface.co/datasets/tantara/Nemotron-Personas-Korea-Qwen3-0.6B-embeddings.clean-gsm8k-aug-qwen3-0.6b
Clean GSM8K-Aug Qwen3-0.6B
This dataset replaces the reasoning steps and answers in
cs-giung/clean-gsm8k-aug
with responses generated by
Qwen/Qwen3-0.6B.
The questions and split ordering match source revision
60f9c039ae300041b5dca5dc1482c8ff1ef1eb47. The model revision is
c1899de289a04d12100db370d81485cdf75e47ca.
Generation
Qwen3 thinking mode was enabled. Sampling used:
temperature: 0.6
top-p: 0.95
top-k: 20
min-p: 0.0
maximum new tokens: 32768
initial seed: 0… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/clean-gsm8k-aug-qwen3-0.6b.AlphaTrade-0.6B-SFT-v0.1-Evaluated-Dataset-2agenttune-apigen-SFT-qwen3-0.6b-tracesAlphaTrade-0.6B-SFT-v0.1-Raw-Dataset-2
