datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flint-section-aware-gemma-4-12b-it
flint-section-aware-gemma12b-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the section-aware-gemma12b variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by unsloth/gemma-4-12b-it (segmenter: unsloth/gemma-4-12b-it), policy
policy/1.1, template caveman_convert/2.0.
Cross-family replication: the section-aware recipe run end-to-end on unsloth/gemma-4-12b-it (self-generated traces, self-voice segmentation and… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-gemma-4-12b-it.crucible-sft-gemma-4-12b-it-mini
crucible-sft-gemma-4-12b-it-mini
Self-distilled SFT dataset of verified reasoning traces from unsloth/gemma-4-12b-it,
built by the reasoning-compression
crucible pipeline: k-sample generation on a decontaminated prompt pool, inline
verification (symbolic math / sandboxed code tests), difficulty banding via
solve rate, and loop-detector filtering on the chosen trace.
Each row: prompt, reasoning (a verified-correct thinking trace when the
domain is verifiable), response, domain… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/crucible-sft-gemma-4-12b-it-mini.
