datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.tibetan-metadata-llm-sft-full
Tibetan metadata LLM SFT dataset (full)
Complete supervised fine-tuning JSONL for title and author span extraction from BDRC outliner segments, built for TiLamb-7B.
Pilot subset
See ganga4364/tibetan-metadata-llm-sft for a 10% pilot subsample (same schema, smaller for quick experiments).
Layout
title/{train,val,test}.jsonl # Alpaca format for LLaMA-Factory
title/{train,val,test}_meta.jsonl
author/{train,val,test}.jsonl
author/{train,val… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-llm-sft-full.CHUNKY-tulu3-SFT-25k-attributes-full
SURF Attributes (Full)
Complete dataset for SURF research and extension.
Paper: Chunky Post-Training
Quick Start
For running SURF, use the minimal dataset: seoirsem/CHUNKY-tulu3-SFT-25k-attributes
uv run -m surf.cli.main sweep \
--attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Dataset Fields
prompt: The query text
response: The model response (if available)
attributes: Raw extracted attributes… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes-full.
