datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.blackline-atlas-training-corpus-v1
Blackline Atlas Training Corpus v1
Blackline Atlas Training Corpus v1 is a license-aware, normalized dataset for
building structured satellite-imagery triage systems over civilian disruption
scenarios. It combines text-only planning examples, paired-image
vision-language examples, hard negatives, and audit/provenance notes into a
training-ready corpus.
The repository is not a raw mirror of third-party datasets. Rows are normalized
for supervised fine-tuning, eval, and reproducible… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/blackline-atlas-training-corpus-v1.codette_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.STEM_train_cot
STEM Image Chain-of-Thought Edit Analysis Dataset
This dataset contains AI-generated Chain-of-Thought (CoT) reasoning for STEM image editing tasks, providing step-by-step analysis of edit operations.
Dataset Structure
The dataset is organized in batches:
Total batches: 26
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
Source image and caption (from previous stage)
Edit command (from original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_cot.beecare-vision-public-bilingual-train
BeeCare Vision Public Bilingual Train
Default train split for Unsloth vision fine-tuning. Each row has image, messages, and metadata.
Use in Unsloth Studio as: yahelr1/beecare-vision-public-bilingual-train.
generative-rlhf-training-data
Generative RLHF-V Online Interaction Training Data
This repository contains the exact 4,053-example multimodal preference mixture
used by our online interaction training runs. The train split preserves the
original training order through training_index.
Composition
Source
Examples
PKU-Alignment/align-anything
2,048
PKU-Alignment/BeaverTails-V
2,005
Total
4,053
The Align-Anything subset was sampled from
text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.Allegator-train
Alleviating Attention Bias for Visual-Informed Text Generation
Training Dataset for Allegator finetuning.
The dataset consists of a subset of LLaVA-Instruct-150k and a subset of Flickr30k, with 99,883 and 31,783 samples from each, respectively.
In detail, LLaVA-Instruct-150K contains 158k language-image instruction following samples, including 58k conversations, 23k descriptions, and 77k complex, 182 reasoning.
We augment LLaVA-Instruct-150K with Flickr30K, which… See the full description on the dataset page: https://huggingface.co/datasets/miso-choi/Allegator-train.STEM_train_edit_predictions_test
STEM Image Edit Predictions Dataset
This dataset contains AI-generated edit predictions for STEM images based on captions and edit commands.
Dataset Structure
The dataset is organized in batches:
Total batches: 1
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
layout_summary: Concise description of source image layout and key elements
edit_analysis: Specific elements to be edited and changes to be made… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_edit_predictions_test.
