CoolFace
16 results

sft-training

leaderonehit /DRT-SFT-8B-training-data DRT-SFT-8B Training Data Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT This dataset contains the SFT training parquet shards used for DRT-SFT-8B. Contents 20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet Total rows: 194,719 Columns: problem_id, content, role, image Downloaded size: about 30.4 GiB Notes The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.textvisual-question-answering100K<n<1M0 likes2.5k downloads3d agoHugging FaceNuTonic /sat-vl-sft-training-ready-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.imagetext-generation100K<n<1M2 likes1.3k downloads5mo agoHugging Facedi-zhang-fdu /Llama-Nemotron-Post-Training-Dataset-SFT-CoT-Only0 likes1k downloads1y agoHugging Facesurrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-data0 likes289 downloads21d agoHugging Facemodel-organisms-for-real /olmo2_1b_sft_checkpoint_oracle_v1-training-data0 likes261 downloads5mo agoHugging FaceLumiOpen /Llama-Nemotron-Post-Training-Dataset-SFT-math-FI Llama-Nemotron-Post-Training-Dataset-SFT-math-FI This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset. The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model. Translation Process The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.texttext-generation1M<n<10M1 likes229 downloads2mo agoHugging Face