datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.agentbench-sft-trajectories-v4-longexp-fullLuauDev-instructions-SFT-full
LuauDev-SFT-FULL 🚀
THIS IS THE FULL VARIANT OF LUAUDEV.
This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models.
These are the models which were used for data generation:
(no specific order)
DiffusionGemma 26B A4B
Deepseek v4 Flash 0731
Deepseek v4.1 flash
Deepseek v4 pro
Nemotron 3 Ultra 550B A55B
dots3 note prev
GPT OSS 120b
Muse Glimmer 30B
GPT OSS 20b
Ling 3.0 flash
And more...
each row has exactly 5 assistant and user… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-full.tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2
ToT CWQ Plan SFT - outputs34_rule_full_pw4_expand_labels_v2
Merged SFT output from local run outputs34_rule_full_pw4_expand_labels_v2.
Version ID
local output dir: tot/sft/outputs34_rule_full_pw4_expand_labels_v2
file: cwq_train_plan.no_mid.jsonl
dataset: CWQ
grouping backend: TOT_REL_GROUPING_BACKEND=rules
parallel workers: 4
strict expand parity: enabled
nested expand labels: enabled
Main difference from earlier runs
This version renders nested Expand… See the full description on the dataset page: https://huggingface.co/datasets/YF0808/tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2.icd_naive_sft_mimic4_fullmath-sft-full-2math-sft-full-llama-3-70b-instructmath-sft-full-8math-sft-full-command-r-2024-03math-sft-full-5icd_naive_sft_mimic3_fullicd_naive_sft_mimic4_fullscience-sft-full-4Full_SFT_plus_Cascade-SFT-Stage-1Full data from SFT merged and Cascade-stage-1
