CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuTonic /sat-image-boundingbox-sft-full NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.imageimage-text-to-text100K<n<1M14 likes5.2k downloads5mo agoHugging Face02NuTonic /sat-vl-sft-training-ready-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.imagetext-generation100K<n<1M2 likes1.3k downloads5mo agoHugging Face03satpalsr /hallucinations-dpotext1K<n<10K5 likes908 downloads3y agoHugging Face04AI45Research /SATraj-OS SATraj-OS: Scaling Agent Trajectories for OSWorld 中文 &nbsp;|&nbsp; English SATraj-OS is a large-scale multimodal graphical user interface (GUI) interaction trajectory dataset for computer-using agents (CUAs), designed for general capability learning and safety training. 📈 Model Performance After joint training on Capability and Safety-v2, the SCOPE model series achieves a stronger balance between general capability on OSWorld and safety on OS-Blind. The x-axis below… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/SATraj-OS.imageimage-text-to-text10K<n<100K4 likes814 downloads1mo agoHugging Face05NuTonic /sat-vl-sft-postprocessed-merged-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.imagetext-generation100K<n<1M0 likes482 downloads5mo agoHugging Face06NuTonic /sat-image-boundingbox-sft NU-TONIC raw SFT init Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py. Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft.imageimage-text-to-text10K<n<100K0 likes426 downloads5mo agoHugging Face07SatsukiVie /FusionAudiotextquestion-answering1M<n<10M9 likes365 downloads1y agoHugging Face08NuTonic /sat-bbox-metadata-sft-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.imagetext-generation100K<n<1M4 likes330 downloads5mo agoHugging Face09Shoozes /LFM-Orbit-SatData LFM Orbit SatData Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon. The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema. Configs Config File Purpose default training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.image1K<n<10K0 likes303 downloads5mo agoHugging Face10LLM4Code /SATBench SATBench: Benchmarking LLMs’ Logical Reasoning via Automated Puzzle Generation from SAT Formulas Paper: https://arxiv.org/abs/2505.14615 Dataset Summary Size: 2,100 puzzles Format: JSONL Data Fields Each JSON object has the following fields: Field Name Description dims List of integers describing the dimensional structure of variables num_vars Total number of variables num_clauses Total number of clauses in the CNF formula readable Readable… See the full description on the dataset page: https://huggingface.co/datasets/LLM4Code/SATBench.tabular1K<n<10K4 likes277 downloads9mo agoHugging Face11PocketDoc /Dans-Logicmaxx-SAT-APtext1K<n<10K0 likes262 downloads2y agoHugging Face12ojayy /logical-sata LOGICAL-SATA LOGICAL-SATA is a reading-comprehension benchmark for compound answer reasoning. Each instance has a paragraph, a question, and four candidate answer options. Every option joins two atomic answers under an explicit logical operator, AND, OR, or NEITHER/NOR. Exactly one option is valid per instance. The dataset is built to isolate logical composition from comprehension, so a model can get every atomic judgment right and still fail if it can't combine them correctly… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-sata.tabularquestion-answering1K<n<10K0 likes204 downloads1mo agoHugging Face13satgeze /Qwen3.5-0.8B-selfdistill Qwen3.5-0.8B self-distillation pairs (DSpark drafter training data) The exact training data behind satgeze/Qwen3.5-0.8B-DSpark: public prompts answered by the target model itself, so a drafter trains on precisely the distribution it will draft for at inference. The unique part is the responses; they were generated by Qwen3.5-0.8B and exist nowhere upstream. Provenance, exactly part source license prompts mlabonne/open-perfectblend Apache-2.0 responses… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.5-0.8B-selfdistill.texttext-generation10K<n<100K1 likes121 downloads2mo agoHugging Face14tasksource /natural-language-satisfiability@misc{https://doi.org/10.48550/arxiv.2211.05417, doi = {10.48550/ARXIV.2211.05417}, url = {https://arxiv.org/abs/2211.05417}, author = {Schlegel, Viktor and Pavlov, Kamen V. and Pratt-Hartmann, Ian}, keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences}, title = {Can Transformers Reason in Fragments of Natural Language?}, publisher = {arXiv}, year = {2022}, copyright =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/natural-language-satisfiability.tabulartext-classification1K<n<10K1 likes103 downloads2y agoHugging Face15sattsansar /Physibench PhysiBench A physics-grounded visual reasoning benchmark testing whether AI vision models can judge the physical plausibility of projectile motion. Inspired by the labeling rigor of ImageNet and the spatial-intelligence focus of Fei-Fei Li's World Labs — PhysiBench targets a narrower, more measurable question: can current vision-language models tell when a trajectory violates real physics? Dataset 300 labeled items, each a rendered PNG image of an object's… See the full description on the dataset page: https://huggingface.co/datasets/sattsansar/Physibench.imagen<1K0 likes96 downloads17d agoHugging Face16hausmer /ukr-tg-satire Ukrainian Telegram satire & troll posts A corpus of 92k posts from 17 public Telegram channels in the satire/irony/parody register, harvested July 2026 via the public Telegram API (history reaching back to 2018 for some channels). The channels are Ukrainian-audience but the posts are a Ukrainian/Russian mix (the main truha and sria_news feeds write mostly in Russian, the regional truexa* branches mostly in Ukrainian), so the corpus carries language: [ru, uk]. Two flavors are… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire.tabulartext-generation10K<n<100K0 likes73 downloads7d agoHugging Face17ChrisRPL /satellite-civilian-conflict-disruption-reporter-v1 Satellite Civilian Conflict Disruption Reporter v1 Dataset ID: ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1 Status This is a valid diagnostic reporter-schema dataset, not the current Blackline Atlas canonical model gate. The canonical compact calibration/gold dataset remains ChrisRPL/satellite-disruption-triage-aux-v2-2. Use this dataset for future schema-simplification experiments only after respecting the mixed source licenses. Do not treat the associated… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1.imageimage-to-textn<1K0 likes66 downloads5mo agoHugging Face18MCINext /synthetic-persian-chatbot-satisfaction-level-classification Dataset Summary Synthetic Persian Chatbot Satisfaction Level Classification (SynPerChatbotSatisfactionLevelClassification) is a Persian (Farsi) dataset designed for the Classification task, specifically to identify the user's satisfaction level in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using the GPT-4o-mini Large Language Model and derived from the Synthetic Persian Chatbot Dataset. Each… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-satisfaction-level-classification.text10K<n<100K0 likes62 downloads1y agoHugging Face19Satyam970 /legal-ai-rag-datatext100K<n<1M0 likes58 downloads6d agoHugging Face20Satori-reasoning /Satori-SWE-two-stage-SFT-datatext10K<n<100K2 likes53 downloads1y agoHugging Face21MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction Dataset Summary Synthetic Persian Chatbot Conversational Sentiment Analysis – Satisfaction is a Persian (Farsi) dataset developed for the Classification task, specifically focused on detecting the emotion of satisfaction in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using the GPT-4o-mini language model. Language(s): Persian (Farsi) Task(s): Classification (Emotion Detection – Satisfaction) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction.text1K<n<10K0 likes52 downloads1y agoHugging Face22youngermax /digital-sat-words-in-context-llmtextn<1K0 likes51 downloads1y agoHugging Face23hausmer /ukr-tg-satire-2 TG channels wave 2 A second scrape wave of 6 public Telegram channels exported 2026-09-15 from their full histories via a read-only MTProto API. Same schema as the main corpora (hausmer/ukr-tg-media, hausmer/ukr-tg-satire). 24,743 text posts across: Foma_memes — Мемарня Sa-chan1917|ИзюмТГ (1,892 posts) zelenskyi_vladimir — Владимир Зеленский (пародия) (612 posts) karikaturnaya_satira — Карикатурная сатира (138 posts) memarnya_rezerv — Ініціативна група «20 см» (1,169 posts)… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire-2.texttext-generation10K<n<100K0 likes51 downloads7d agoHugging Face24satpalsr /yes-no-factstextn<1K1 likes49 downloads2y agoHugging Face25SigmaPrep /sat-math-formula-sheet SAT Math Formula Sheet The 24 formulas and concepts the SAT does not give you on its reference sheet. The digital SAT provides a reference sheet on every math question with area, circumference, volume and right-triangle formulas. It does not provide slope, the quadratic forms, the discriminant, exponent rules, percent change, exponential growth, probability, or the circle equation. These 24 cover what it leaves out. Contents formulas.json holds 24 records:… See the full description on the dataset page: https://huggingface.co/datasets/SigmaPrep/sat-math-formula-sheet.textn<1K0 likes48 downloads11d agoHugging Face26hausmer /truha-news-satire Truha news satire (Ukrainian) A synthetic, LLM-generated dataset of Ukrainian "news" written in the truha satirical register — short, meta-ironic, punchline-driven fake-news items. The corpus was produced by distilling a target style (a Ukrainian satirical news persona) into generated examples; no real user data is included. Intended use: style-transfer / imitation training for a Ukrainian satirical news-writing assistant. Each example is a (system, instruction, output) triple.… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truha-news-satire.texttext-generation10K<n<100K0 likes44 downloads8d agoHugging Face27mishhkaa /satellite-water-detectiontabularn<1K0 likes42 downloads10mo agoHugging Face28satgeze /Qwen3.6-27B-selfdistill Qwen3.6-27B self-distillation pairs (DSpark drafter training data) The exact training data behind satgeze/Qwen3.6-27B-DSpark (2.5-2.7x measured decode speedup in llama.cpp): public prompts answered by the target model itself, so the drafter trains on precisely the distribution it drafts for at inference. The responses were generated by Qwen3.6-27B and exist nowhere upstream. Provenance, exactly part source license prompts mlabonne/open-perfectblend (12K… See the full description on the dataset page: https://huggingface.co/datasets/satgeze/Qwen3.6-27B-selfdistill.texttext-generation10K<n<100K0 likes42 downloads2mo agoHugging Face29sata-bench /sata-bench Cite @misc{xu2025satabenchselectapplybenchmark, title={SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions}, author={Weijie Xu and Shixian Cui and Xi Fang and Chi Xue and Stephanie Eckman and Chandan Reddy}, year={2025}, eprint={2506.00643}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2506.00643}, } Select-All-That-Apply Benchmark (SATA-bench) Dataset Desciption SATA-Bench is… See the full description on the dataset page: https://huggingface.co/datasets/sata-bench/sata-bench.textquestion-answering1K<n<10K9 likes41 downloads1y agoHugging Face30ChristianZ97 /putnambench-satp PutnamBench-SATP — Lean 4 (672 problems, SATP-normalized) PutnamBench Lean 4 problems normalized for the SATP-DSP-Eval pipeline. Upstream: trishullab/PutnamBench (lean4/src/*.lean + informal/putnam.json, parsed via this repo's datasets/putnam_bench/parse_lean4.py). Differences from upstream Field Upstream This repo formal_statement ends with := sorry normalized to := by (no trailing tactic) uuid absent sha256(canonical(formal_statement))[:16] after… See the full description on the dataset page: https://huggingface.co/datasets/ChristianZ97/putnambench-satp.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.