datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.sft_training_corpusk3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.colab-training-demo-sft
colab-training demo SFT dataset
500 synthetic two-digit addition pairs in messages (chat) format.
Generated for validating the colab_training QLoRA pipeline; after training,
ask the adapter "What is 34 + 58?" and expect "34 + 58 = 92".
Astral-1.5-Post-Training-Dataset-SFT
Astral 1.5 Post-Training Dataset
A albeit smaller, yet higher-quality reasoning dataset combining mathematics, code, and general stem used in the training of the Astral 1.5 model family.
Dataset Description
This dataset merges four datasets to create a high quality 25 thousand example dataset. With the size of the dataset, we rely on the principle that quality > quantity leads to better model performance.
Dataset Composition
Setup
General STEM:… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/Astral-1.5-Post-Training-Dataset-SFT.chess_sft_training_datasft_training_dataABot-PhysWorld_SFT_Training_Data_v1_OXETFRank-sft-training-dataTraining_TAO_V2_SFTSFT_training_datacombined-sft-training-data-v20250824_MiroSystemPrompt
