CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuTonic /sat-image-boundingbox-sft-full NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.imageimage-text-to-text100K<n<1M14 likes5.2k downloads5mo agoHugging Face02hamishivi /tmax-sft-full-20260403text100K<n<1M0 likes274 downloads5mo agoHugging Face03mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes189 downloads7mo agoHugging Face04hamishivi /tmax-sft-full-20260317text100K<n<1M0 likes189 downloads5mo agoHugging Face05osieosie /tmax-sft-full-20260310text100K<n<1M0 likes178 downloads7mo agoHugging Face06osieosie /tmax-sft-full-20260403text100K<n<1M0 likes116 downloads6mo agoHugging Face07seoirsem /CHUNKY-tulu3-SFT-25k-attributes-full SURF Attributes (Full) Complete dataset for SURF research and extension. Paper: Chunky Post-Training Quick Start For running SURF, use the minimal dataset: seoirsem/CHUNKY-tulu3-SFT-25k-attributes uv run -m surf.cli.main sweep \ --attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \ --rubric rubrics/rebuttal.yaml \ -o results/ Dataset Fields prompt: The query text response: The model response (if available) attributes: Raw extracted attributes… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes-full.texttext-generation100K<n<1M0 likes113 downloads8mo agoHugging Face08osieosie /tmax-sft-full-20260315text100K<n<1M0 likes103 downloads6mo agoHugging Face09osieosie /tmax-sft-full-20260317text100K<n<1M0 likes91 downloads6mo agoHugging Face10osieosie /tmax-sft-full-20260513text100K<n<1M0 likes80 downloads4mo agoHugging Face11osieosie /tmax-sft-full-20260513-bad-tool-call-filteredtext100K<n<1M0 likes50 downloads4mo agoHugging Face12yosubshin /m2sv-sft-11k-fullimage10K<n<100K0 likes47 downloads11mo agoHugging Face13zsqzz /sft-mathhard-medium-with-thinking-full-paralleltabular1K<n<10K0 likes45 downloads1y agoHugging Face14Gragroo /therapist-sft-full_traintext10K<n<100K0 likes44 downloads2y agoHugging Face15seoirsem /tulu3-SFT-500k-25k-data-attributes-full SURF Attributes (Full) Complete dataset for SURF research and extension. Quick Start For running SURF, use the minimal dataset: seoirsem/tulu3-SFT-500k-25k-data-attributes uv run -m surf.cli.main sweep \ --attributes seoirsem/tulu3-SFT-500k-25k-data-attributes \ --rubric rubrics/rebuttal.yaml \ -o results/ Dataset Fields prompt: The query text response: The model response (if available) attributes: Raw extracted attributes (10 per query)… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/tulu3-SFT-500k-25k-data-attributes-full.text100K<n<1M0 likes38 downloads8mo agoHugging Face16Gragroo /therapist-sft-formatted-full_traintext10K<n<100K1 likes35 downloads2y agoHugging Face17Gragroo /therapist-sft-full_train-validationtext10K<n<100K1 likes35 downloads2y agoHugging Face18GENIAC-Team-Ozaki /tuninig-dataset_pref_20pct_v2_full-sft-finetuned-stage4-iter86000-v2text10K<n<100K0 likes28 downloads2y agoHugging Face19GENIAC-Team-Ozaki /tuninig-dataset_pref_20pct_v3_full-sft-finetuned-stage4-iter86000-v3text10K<n<100K0 likes27 downloads2y agoHugging Face20SiliangZ /full_ultrachat_200k_vs_sft_with_spin_iter0 Dataset Card for "full_ultrachat_200k_vs_sft_with_spin_iter0" More Information needed text100K<n<1M0 likes27 downloads2y agoHugging Face21lihaoxin2020 /qwen-instruct-synthetic_1_stem_only-sft-full-supergpqa-r1text1K<n<10K0 likes27 downloads1y agoHugging Face22sabia0080 /agentbench-sft-trajectories-v4-longexp-fulltext1K<n<10K0 likes27 downloads7mo agoHugging Face23GENIAC-Team-Ozaki /tuninig-dataset_pref_20pct_full-sft-finetuned-stage4-iter86000text10K<n<100K0 likes26 downloads2y agoHugging Face24Pinkstack /LuauDev-instructions-SFT-full LuauDev-SFT-FULL 🚀 THIS IS THE FULL VARIANT OF LUAUDEV. This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models. These are the models which were used for data generation: (no specific order) DiffusionGemma 26B A4B Deepseek v4 Flash 0731 Deepseek v4.1 flash Deepseek v4 pro Nemotron 3 Ultra 550B A55B dots3 note prev GPT OSS 120b Muse Glimmer 30B GPT OSS 20b Ling 3.0 flash And more... each row has exactly 5 assistant and user… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-full.text10K<n<100K1 likes26 downloads3d agoHugging Face25selfcorrexp2 /llama31_sft_non_delete_fulltabular100K<n<1M0 likes25 downloads2y agoHugging Face26ChavyvAkvar /SFT-GRPO-dataset-v2-full Dataset Card for "SFT-GRPO-dataset-v2-full" More Information needed tabular100K<n<1M0 likes24 downloads1y agoHugging Face27Tandogan /sft_dataset_fulltext10K<n<100K0 likes23 downloads1y agoHugging Face28raca-workspace-v1 /algorithmic-sft-full-eval-v4 algorithmic-sft-full-eval-v4 Aggregate eval results: 10 models x 4 domains x 3 splits with bootstrap 95% CIs Dataset Info Rows: 42 Columns: 8 Columns Column Type Description model Value('string') HuggingFace model ID (LoRA adapter name) domain Value('string') Task domain: formal_logic, conlang_morphology, cellular_automata, long_arithmetic type Value('string') Training type: algo (algorithmic SFT) or distill (QwQ distillation) split… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algorithmic-sft-full-eval-v4.tabularn<1K0 likes20 downloads6mo agoHugging Face29artnoage /sft_fulltext100K<n<1M0 likes18 downloads2y agoHugging Face30TAUR-dev /C-SFT_OT_Partial_Nov17_p2_reflections5_formats-C_fulltext10K<n<100K0 likes17 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.