CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuTonic /sat-vl-sft-training-ready-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.imagetext-generation100K<n<1M2 likes1.3k downloads5mo agoHugging Face02TharunSivamani /sft_training_corpustext1M<n<10M0 likes142 downloads9mo agoHugging Face03Training-Datasmith /k3-sft-cc0-flan Dataset Card for K3 SFT CC0 FLAN 844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort views; adaptive is the recommended default for quality-conscious SFT mixing. Dataset Details Curated by: Training Datasmith Teacher: kimi-k3 via deltafin (local inference) Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.tabulartext-classification1K<n<10K0 likes103 downloads5d agoHugging Face04pnutnam /colab-training-demo-sft colab-training demo SFT dataset 500 synthetic two-digit addition pairs in messages (chat) format. Generated for validating the colab_training QLoRA pipeline; after training, ask the adapter "What is 34 + 58?" and expect "34 + 58 = 92". texttext-generationn<1K0 likes39 downloads19d agoHugging Face05LucidityAI /Astral-1.5-Post-Training-Dataset-SFT Astral 1.5 Post-Training Dataset A albeit smaller, yet higher-quality reasoning dataset combining mathematics, code, and general stem used in the training of the Astral 1.5 model family. Dataset Description This dataset merges four datasets to create a high quality 25 thousand example dataset. With the size of the dataset, we rely on the principle that quality > quantity leads to better model performance. Dataset Composition Setup General STEM:… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/Astral-1.5-Post-Training-Dataset-SFT.text10K<n<100K0 likes16 downloads10mo agoHugging Face06czovekboti /chess_sft_training_datatext10K<n<100K0 likes14 downloads11mo agoHugging Face07huyhuung /sft_training_datatext10K<n<100K0 likes13 downloads1y agoHugging Face08ryanhoangt /ABot-PhysWorld_SFT_Training_Data_v1_OXEtext100K<n<1M0 likes11 downloads4mo agoHugging Face09Johnnyfans /TFRank-sft-training-datatext100K<n<1M2 likes5 downloads1y agoHugging Face10telconemotronv3 /Training_TAO_V2_SFTtext100K<n<1M0 likes5 downloads8mo agoHugging Face11Ramkumar-AI-developer /SFT_training_datatext10K<n<100K0 likes5 downloads3mo agoHugging Face12rl-rag /combined-sft-training-data-v20250824_MiroSystemPrompttext1K<n<10K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.