datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-vl-sft-postprocessed-merged-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.Sumtablets_Merged
SumTablets Merged Multimodal Dataset
A merged, deduplicated, quality-filtered training corpus combining cuneiform tablet text records with tablet photography and lineart, structured for vision-language model (VLM) fine-tuning.
Target model: Qwen3-VL-8B-Instruct via Unsloth StudioCombined license: CC-BY-4.0 (most restrictive of the two source licenses applies)
Quick Start
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/TRACCERR/Sumtablets_Merged.finance-legal-mrc_merged-table
데이터셋 설명
shchoice/finance-legal-mrc 데이터 중 병합된 테이블만 추출한 뒤 이미지와 함께 저장한 데이터입니다.
