datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.hallucinations-dpoSATraj-OS
SATraj-OS: Scaling Agent Trajectories for OSWorld
中文 | English
SATraj-OS is a large-scale multimodal graphical user interface (GUI) interaction trajectory dataset for computer-using agents (CUAs), designed for general capability learning and safety training.
📈 Model Performance
After joint training on Capability and Safety-v2, the SCOPE model series achieves a stronger balance between general capability on OSWorld and safety on OS-Blind. The x-axis below… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/SATraj-OS.sat-vl-sft-postprocessed-merged-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.sat-image-boundingbox-sft
NU-TONIC raw SFT init
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py.
Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft.FusionAudiosat-bbox-metadata-sft-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.LFM-Orbit-SatData
LFM Orbit SatData
Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon.
The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema.
Configs
Config
File
Purpose
default
training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.SATBench
SATBench: Benchmarking LLMs’ Logical Reasoning via Automated Puzzle Generation from SAT Formulas
Paper: https://arxiv.org/abs/2505.14615
Dataset Summary
Size: 2,100 puzzles
Format: JSONL
Data Fields
Each JSON object has the following fields:
Field Name
Description
dims
List of integers describing the dimensional structure of variables
num_vars
Total number of variables
num_clauses
Total number of clauses in the CNF formula
readable
Readable… See the full description on the dataset page: https://huggingface.co/datasets/LLM4Code/SATBench.Dans-Logicmaxx-SAT-APlogical-sata
LOGICAL-SATA
LOGICAL-SATA is a reading-comprehension benchmark for compound answer reasoning. Each instance has a paragraph, a question, and four candidate answer options. Every option joins two atomic answers under an explicit logical operator, AND, OR, or NEITHER/NOR. Exactly one option is valid per instance. The dataset is built to isolate logical composition from comprehension, so a model can get every atomic judgment right and still fail if it can't combine them correctly… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-sata.Qwen3.5-0.8B-selfdistill
Qwen3.5-0.8B self-distillation pairs (DSpark drafter training data)
The exact training data behind
satgeze/Qwen3.5-0.8B-DSpark: public
prompts answered by the target model itself, so a drafter trains on precisely the distribution
it will draft for at inference. The unique part is the responses; they were generated by
Qwen3.5-0.8B and exist nowhere upstream.
Provenance, exactly
part
source
license
prompts
mlabonne/open-perfectblend
Apache-2.0
responses… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.5-0.8B-selfdistill.natural-language-satisfiability@misc{https://doi.org/10.48550/arxiv.2211.05417,
doi = {10.48550/ARXIV.2211.05417},
url = {https://arxiv.org/abs/2211.05417},
author = {Schlegel, Viktor and Pavlov, Kamen V. and Pratt-Hartmann, Ian},
keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
title = {Can Transformers Reason in Fragments of Natural Language?},
publisher = {arXiv},
year = {2022},
copyright =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/natural-language-satisfiability.Physibench
PhysiBench
A physics-grounded visual reasoning benchmark testing whether AI vision models can judge
the physical plausibility of projectile motion.
Inspired by the labeling rigor of ImageNet and the spatial-intelligence focus of
Fei-Fei Li's World Labs — PhysiBench targets a narrower, more measurable question:
can current vision-language models tell when a trajectory violates real physics?
Dataset
300 labeled items, each a rendered PNG image of an object's… See the full description on the dataset page: https://huggingface.co/datasets/sattsansar/Physibench.ukr-tg-satire
Ukrainian Telegram satire & troll posts
A corpus of 92k posts from 17 public Telegram channels in the
satire/irony/parody register, harvested July 2026 via the public Telegram
API (history reaching back to 2018 for some channels). The channels are
Ukrainian-audience but the posts are a Ukrainian/Russian mix (the
main truha and sria_news feeds write mostly in Russian, the regional
truexa* branches mostly in Ukrainian), so the corpus carries
language: [ru, uk].
Two flavors are… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire.satellite-civilian-conflict-disruption-reporter-v1
Satellite Civilian Conflict Disruption Reporter v1
Dataset ID: ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1
Status
This is a valid diagnostic reporter-schema dataset, not the current Blackline Atlas canonical model gate. The canonical compact calibration/gold dataset remains ChrisRPL/satellite-disruption-triage-aux-v2-2.
Use this dataset for future schema-simplification experiments only after respecting the mixed source licenses. Do not treat the associated… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1.synthetic-persian-chatbot-satisfaction-level-classification
Dataset Summary
Synthetic Persian Chatbot Satisfaction Level Classification (SynPerChatbotSatisfactionLevelClassification) is a Persian (Farsi) dataset designed for the Classification task, specifically to identify the user's satisfaction level in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using the GPT-4o-mini Large Language Model and derived from the Synthetic Persian Chatbot Dataset. Each… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-satisfaction-level-classification.legal-ai-rag-dataSatori-SWE-two-stage-SFT-datasynthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction
Dataset Summary
Synthetic Persian Chatbot Conversational Sentiment Analysis – Satisfaction is a Persian (Farsi) dataset developed for the Classification task, specifically focused on detecting the emotion of satisfaction in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using the GPT-4o-mini language model.
Language(s): Persian (Farsi)
Task(s): Classification (Emotion Detection – Satisfaction)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction.digital-sat-words-in-context-llmukr-tg-satire-2
TG channels wave 2
A second scrape wave of 6 public Telegram channels
exported 2026-09-15 from their full histories via a
read-only MTProto API. Same schema as the main corpora
(hausmer/ukr-tg-media,
hausmer/ukr-tg-satire).
24,743 text posts across:
Foma_memes — Мемарня Sa-chan1917|ИзюмТГ (1,892 posts)
zelenskyi_vladimir — Владимир Зеленский (пародия) (612 posts)
karikaturnaya_satira — Карикатурная сатира (138 posts)
memarnya_rezerv — Ініціативна група «20 см» (1,169 posts)… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire-2.yes-no-factssat-math-formula-sheet
SAT Math Formula Sheet
The 24 formulas and concepts the SAT does not give you on its reference sheet.
The digital SAT provides a reference sheet on every math question with area,
circumference, volume and right-triangle formulas. It does not provide slope,
the quadratic forms, the discriminant, exponent rules, percent change,
exponential growth, probability, or the circle equation. These 24 cover what it
leaves out.
Contents
formulas.json holds 24 records:… See the full description on the dataset page: https://huggingface.co/datasets/SigmaPrep/sat-math-formula-sheet.truha-news-satire
Truha news satire (Ukrainian)
A synthetic, LLM-generated dataset of Ukrainian "news" written in the
truha satirical register — short, meta-ironic, punchline-driven fake-news
items. The corpus was produced by distilling a target style (a Ukrainian
satirical news persona) into generated examples; no real user data is
included.
Intended use: style-transfer / imitation training for a Ukrainian satirical
news-writing assistant. Each example is a (system, instruction, output) triple.… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truha-news-satire.satellite-water-detectionQwen3.6-27B-selfdistill
Qwen3.6-27B self-distillation pairs (DSpark drafter training data)
The exact training data behind
satgeze/Qwen3.6-27B-DSpark (2.5-2.7x
measured decode speedup in llama.cpp): public prompts answered by the target model itself, so
the drafter trains on precisely the distribution it drafts for at inference. The responses were
generated by Qwen3.6-27B and exist nowhere upstream.
Provenance, exactly
part
source
license
prompts
mlabonne/open-perfectblend (12K… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.6-27B-selfdistill.sata-bench
Cite
@misc{xu2025satabenchselectapplybenchmark,
title={SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions},
author={Weijie Xu and Shixian Cui and Xi Fang and Chi Xue and Stephanie Eckman and Chandan Reddy},
year={2025},
eprint={2506.00643},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.00643},
}
Select-All-That-Apply Benchmark (SATA-bench) Dataset Desciption
SATA-Bench is… See the full description on the dataset page: https://huggingface.co/datasets/sata-bench/sata-bench.putnambench-satp
PutnamBench-SATP — Lean 4 (672 problems, SATP-normalized)
PutnamBench Lean 4 problems normalized for the SATP-DSP-Eval pipeline.
Upstream: trishullab/PutnamBench
(lean4/src/*.lean + informal/putnam.json, parsed via this repo's
datasets/putnam_bench/parse_lean4.py).
Differences from upstream
Field
Upstream
This repo
formal_statement
ends with := sorry
normalized to := by (no trailing tactic)
uuid
absent
sha256(canonical(formal_statement))[:16] after… See the full description on the dataset page: https://huggingface.co/datasets/ChristianZ97/putnambench-satp.
