datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.tmax-sft-full-20260403full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.tmax-sft-full-20260317tmax-sft-full-20260310tmax-sft-full-20260403CHUNKY-tulu3-SFT-25k-attributes-full
SURF Attributes (Full)
Complete dataset for SURF research and extension.
Paper: Chunky Post-Training
Quick Start
For running SURF, use the minimal dataset: seoirsem/CHUNKY-tulu3-SFT-25k-attributes
uv run -m surf.cli.main sweep \
--attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Dataset Fields
prompt: The query text
response: The model response (if available)
attributes: Raw extracted attributes… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes-full.tmax-sft-full-20260315tmax-sft-full-20260317tmax-sft-full-20260513tmax-sft-full-20260513-bad-tool-call-filteredm2sv-sft-11k-fullsft-mathhard-medium-with-thinking-full-paralleltherapist-sft-full_traintulu3-SFT-500k-25k-data-attributes-full
SURF Attributes (Full)
Complete dataset for SURF research and extension.
Quick Start
For running SURF, use the minimal dataset: seoirsem/tulu3-SFT-500k-25k-data-attributes
uv run -m surf.cli.main sweep \
--attributes seoirsem/tulu3-SFT-500k-25k-data-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Dataset Fields
prompt: The query text
response: The model response (if available)
attributes: Raw extracted attributes (10 per query)… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/tulu3-SFT-500k-25k-data-attributes-full.therapist-sft-formatted-full_traintherapist-sft-full_train-validationtuninig-dataset_pref_20pct_v2_full-sft-finetuned-stage4-iter86000-v2tuninig-dataset_pref_20pct_v3_full-sft-finetuned-stage4-iter86000-v3full_ultrachat_200k_vs_sft_with_spin_iter0
Dataset Card for "full_ultrachat_200k_vs_sft_with_spin_iter0"
More Information needed
qwen-instruct-synthetic_1_stem_only-sft-full-supergpqa-r1agentbench-sft-trajectories-v4-longexp-fulltuninig-dataset_pref_20pct_full-sft-finetuned-stage4-iter86000LuauDev-instructions-SFT-full
LuauDev-SFT-FULL 🚀
THIS IS THE FULL VARIANT OF LUAUDEV.
This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models.
These are the models which were used for data generation:
(no specific order)
DiffusionGemma 26B A4B
Deepseek v4 Flash 0731
Deepseek v4.1 flash
Deepseek v4 pro
Nemotron 3 Ultra 550B A55B
dots3 note prev
GPT OSS 120b
Muse Glimmer 30B
GPT OSS 20b
Ling 3.0 flash
And more...
each row has exactly 5 assistant and user… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-full.llama31_sft_non_delete_fullSFT-GRPO-dataset-v2-full
Dataset Card for "SFT-GRPO-dataset-v2-full"
More Information needed
sft_dataset_fullalgorithmic-sft-full-eval-v4
algorithmic-sft-full-eval-v4
Aggregate eval results: 10 models x 4 domains x 3 splits with bootstrap 95% CIs
Dataset Info
Rows: 42
Columns: 8
Columns
Column
Type
Description
model
Value('string')
HuggingFace model ID (LoRA adapter name)
domain
Value('string')
Task domain: formal_logic, conlang_morphology, cellular_automata, long_arithmetic
type
Value('string')
Training type: algo (algorithmic SFT) or distill (QwQ distillation)
split… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algorithmic-sft-full-eval-v4.sft_fullC-SFT_OT_Partial_Nov17_p2_reflections5_formats-C_full
