datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brief-composer-sft-v1
BriefComposer SFT
Multi-image analytical brief rows composed from completed FireWatch, OceanScout, LandShift, and FloodPulse dataset folders (metadata/ + images/). Each sample stitches 1–4 images and metadata-derived headlines into one executive-style assistant reply.
Record counts (this build)
Split
JSONL lines
train
6307
validation
851
test
842
total
8000
Inputs
Source roots: one or more --source-root directories (each must contain… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/brief-composer-sft-v1.docker-compose-20000xSFT dataset with 20k examples of Docker run commands being converted into Docker Compose file format.
Each instance follows this format:
{
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "I need the docker-compose configuration that replicates this run: docker run --pull always registry.k8s.io/lapithae/recoverableness:sha-ec12e54"
},
{
"role": "assistant",
"content":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/docker-compose-20000x.android-kotlin-compose-compiler-verified
Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset
5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs.
Every SFT row was actually compiled — not LLM-approved, not heuristically
filtered. A subset was verified behaviorally by running JUnit tests.
Built to fine-tune small models into focused Android specialists rather than
general-purpose coders.
Why this exists
Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.brief-composer-sft-v1-duplicate
BriefComposer SFT
Multi-image analytical brief rows composed from completed FireWatch, OceanScout, LandShift, and FloodPulse dataset folders (metadata/ + images/). Each sample stitches 1–4 images and metadata-derived headlines into one executive-style assistant reply.
Record counts (this build)
Split
JSONL lines
train
2355
validation
310
test
335
total
3000
Inputs
Source roots: one or more --source-root directories (each must contain… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/brief-composer-sft-v1-duplicate.jazz-composer-datasetWriting_Composer
