datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SSDD
SSDD: SAR Ship Detection Dataset (Object Detection)
Unofficial redistribution of SSDD (SAR Ship Detection Dataset), reformatted into a standardized YOLO-compatible directory layout with a seeded validation split.
Disclaimer
This repository is not an official release of SSDD.
SSDD was created by Tianwen Zhang, Xiaoling Zhang, Jianwei Li, and co-authors, who retain all copyright. This repository does not claim ownership of any images, annotations, or… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/SSDD.ssd-perfectblend-qwen3-4b-regen
open-perfectblend regenerated by Qwen/Qwen3-4B
All 1,349,810 rows of the mlabonne/open-perfectblend train split (95%
split, 53 invalid rows skipped) with every assistant turn regenerated by
Qwen/Qwen3-4B via sglang (temperature 0.7, top-p 0.8, top-k 20, min-p 0,
max 4096 tokens, thinking disabled). Zero generation errors. This file is
the durable input the DSpark activation cache is rebuilt from for the
ssd-dspark-qwen3-4b / ssd-dspark-prog-qwen3-4b drafters.
ssd-perfectblend-qwen3-30b-a3b-regen
open-perfectblend regenerated by Qwen/Qwen3-30B-A3B
All 1,349,810 rows of the mlabonne/open-perfectblend train split (95%
split, 53 invalid rows skipped) with every assistant turn regenerated by
Qwen/Qwen3-30B-A3B via sglang (temperature 0.7, top-p 0.8, top-k 20, min-p 0,
max 4096 tokens, thinking disabled). Zero generation errors; mean 887 new
tokens per row. This file is the durable input the DSpark activation cache
is rebuilt from for the ssd-dspark-prog-qwen3-30b-a3b drafter.
Qwen3.5-9B-SSD
SSD Dataset Replication (Qwen3.5-9B)
This dataset is a replication of the "Embarrassingly Simple Self-Distillation Improves Code Generation" (SSD) paper (arXiv:2604.01193).
Overview
The dataset contains coding problems and their corresponding solutions generated by Qwen3.5-9B using high-temperature sampling (T=1.1) to explore the model's latent capabilities. This approach, known as SSD, focuses on "self-distillation" where a model's own correct but non-greedy outputs are… See the full description on the dataset page: https://huggingface.co/datasets/wrmedford/Qwen3.5-9B-SSD.Gemma-4-E4B-it-SSD
SSD Dataset Replication (Gemma-4-E4B-it)
This dataset is a replication of the "Embarrassingly Simple Self-Distillation Improves Code Generation" (SSD) paper (arXiv:2604.01193).
Overview
The dataset contains coding problems and their corresponding solutions generated by Gemma-4-E4B-it using high-temperature sampling (T=1.1) to explore the model's latent capabilities. This approach, known as SSD, focuses on "self-distillation" where a model's own correct but non-greedy… See the full description on the dataset page: https://huggingface.co/datasets/wrmedford/Gemma-4-E4B-it-SSD.ssd-math-v1p1-e06m-fixed-batch-16
SSD Math V1.1-E06M fixed batch
Tiny 16-record SFT JSONL fixture used for the V1.1-E06M 4B single-batch LR diagnostic.
The records are generated math-reasoning traces from the local v0 overfit fixture artifacts/overfit/v0_unique_batch_16.jsonl.
obsidian-bases-query-v1
Obsidian Bases Query Dataset
Training data for a small language model that converts natural language questions into Obsidian Bases .base file format.
Format
instruction: Natural language question
output: JSON .base file content
Usage
For fine-tuning Qwen 3 0.6B with TRL/SFT.
Stats
1,000 training samples
Generated via Claude Sonnet 4
Based on real vault schema (10 entity types)
newpictureobsidian-bases-query-v2-compactnewsssft_test
SFT Mix
This dataset contains the SFT-style sources from the local phase2 JSONL export:
nemotron-pretraining-sft-code/Nemotron-SFT-Code
nemotron-pretraining-sft-general/Nemotron-SFT-General
nemotron-pretraining-sft-math/Nemotron-SFT-MATH
nemotron-pretraining-stem-sft/Nemotron-Pretraining-STEM-SFT
Local source root:
/mnt/nfs/huangyi/phase2_sample_jsonl
Approximate local size:
408 JSONL files
1.40 TiB apparent data size
The repository also includes:
sources.yaml: source glob… See the full description on the dataset page: https://huggingface.co/datasets/Ssdw1r/sft_test.
