CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuTonic /sat-image-boundingbox-sft-full NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.imageimage-text-to-text100K<n<1M14 likes5.2k downloads5mo agoHugging Face02jm-rt /arvo-vulnsmith-full ARVO CyberGym-format smoke dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It contains 10 ARVO tasks that are outside the original CyberGym set. textn<1K0 likes2k downloads2mo agoHugging Face03KiteFishAI /arxiv-tex-corpus-fullarxiv-tex-corpus-full (80GB) Large-scale LaTeX corpus from arXiv (math, CS, physics, statistics) 📄 Paper: https://arxiv.org/abs/2602.17288 📚 Overview arxiv-tex-corpus-full (80GB) is a large-scale dataset of LaTeX source content extracted from papers hosted on arXiv. This version contains approximately 80GB of structured JSONL data, restricted to the following arXiv categories: math cs hep-th hep-ph quant-ph stat.ML stat.TH The dataset is designed for research in: Large… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-full.texttext-generation100K<n<1M13 likes1.2k downloads7mo agoHugging Face04edev2000 /amc12-full AMC12 Dataset (Research-Oriented) A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks. This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research. 📘 Introduction The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/edev2000/amc12-full.text1K<n<10K3 likes922 downloads27d agoHugging Face05SetFit /yelp_review_fulltext100K<n<1M1 likes921 downloads5y agoHugging Face06osbm /prostate128_t2_anatomy_nnUNet_3d_fullres_20_epochtextn<1K0 likes833 downloads3y agoHugging Face07MoreThought /DeepSWEGym2-Full Dataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.texttext-generation100K<n<1M1 likes825 downloads15d agoHugging Face08osbm /prostate158_nnUNet_3d_fullres_20_epochtextn<1K0 likes713 downloads3y agoHugging Face09marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes712 downloads1mo agoHugging Face10MoreThought /DeepSWEGym-Full Dataset Description This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.texttext-generation10K<n<100K1 likes676 downloads16d agoHugging Face11tonychenxyz /ruler-full RULER Benchmark (Full) Complete RULER benchmark dataset with all 13 tasks across 6 context lengths (4K to 128K tokens). Overview Metric Value Total Samples 78,000 (39,000 per variant) Tasks 13 Context Lengths 4K, 8K, 16K, 32K, 64K, 128K Samples per Config 500 Variants memwrap, plain Tasks Retrieval (NIAH - Needle in a Haystack) niah_single_1, niah_single_2, niah_single_3 - Single needle variants niah_multikey_1… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/ruler-full.text10K<n<100K0 likes553 downloads8mo agoHugging Face12birdsql /livesqlbench-base-full-v1 🚀 LiveSQLBench-Base-Full-v1 A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-full-v1.textn<1K3 likes533 downloads3mo agoHugging Face13mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes467 downloads11mo agoHugging Face14FDAbench2026 /FDAbench-Full v1.1 Update (2026-08-06) — multiple split Strengthened the cross-source requirement that multiple-choice tasks are designed around (selecting all correct options should require integrating both the SQL result and the retrieved documents): 264 of 760 tasks were revised, with task IDs, databases, and gold SQL unchanged. Documents-only accuracy drops from 61.5% to 38.7% while full-evidence accuracy stays at 80.6% (3 frontier models, strict exact set match). Diversified the number… See the full description on the dataset page: https://huggingface.co/datasets/FDAbench2026/FDAbench-Full.text1K<n<10K1 likes459 downloads2mo agoHugging Face15birdsql /bird-interact-full🌐 Website • 📄 Paper (ICLR 2026 Oral) • 💻 GitHub • 🗄️ bird-interact-lite • 🗄️ bird-interact-full • 🗄️ LiveSQLBench 🧸 Overview BIRD-INTERACT, an interactive text-to-SQL benchmark, re-imagines Text-to-SQL evaluation via lens of dynamic interactions which is built on top of single-turn unambiguous T2S tasks from LiveSQLBench. The environment blends a hierarchical knowledge base, database documentation and a function-driven user simulator to recreate authentic enterprise… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-interact-full.textn<1K5 likes404 downloads8mo agoHugging Face16MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes369 downloads6mo agoHugging Face17marin-dna /vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered Review status: draft generated for issue #473 review before upload. Human-anchored 255 bp vertebrate sequences for the ccre_enhancer_centered cohort under the full_window policy. The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.tabular10M<n<100M0 likes363 downloads1mo agoHugging Face18aslawliet /flan2021-full Task Name FLAN-2021 -> 70 { "ag_news_subset": 108497, "ai2_arc/ARC-Challenge": 829, "ai2_arc/ARC-Easy": 1927, "aeslc": 13187, "anli/r1": 15361, "anli/r2": 41133, "anli/r3": 91048, "bool_q": 8343, "cnn_dailymail": 259607, "coqa": 6456, "cosmos_qa": 22996, "definite_pronoun_resolution": 1079, "drop": 70045, "fix_punct": 25690, "gem/common_gen": 60936, "gem/dart": 56724, "gem/e2e_nlg": 30337, "gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.texttext-generation10M<n<100M2 likes330 downloads2y agoHugging Face19VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes280 downloads4mo agoHugging Face20OpenGVLab /InternVid-Full InternVid InternVid-Full We present InternVid-230M, a full set of this dataset, consisting of 230 million video clips, with generated high-quality captions for publicly available web videos. Download The 230M samples are provided in jsonlines file. Columns include the videoID, timestamps, generated caption and their UMT similarity scores. How to Use from datasets import load_dataset dataset = load_dataset("OpenGVLab/InternVid-Full") Method… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVid-Full.text10M<n<100M19 likes275 downloads2y agoHugging Face21DylanRiden /sets_lego_omr_full Dataset Card for sets_lego_omr_full This dataset combines official LEGO sets from LDRAW OMR with metadata from Rebrickable. Each entry contains the full MPD file as a string plus associated metadata such as set number, name, theme, year, and parts count. It is intended for building LLM fine-tuning datasets for LEGO model generation tasks. Dataset Details Dataset Sources OMR files: LDraw OMR Library Metadata: Rebrickable MPD file format: LDraw File Format… See the full description on the dataset page: https://huggingface.co/datasets/DylanRiden/sets_lego_omr_full.texttext-to-3d1K<n<10K2 likes263 downloads8mo agoHugging Face22allenai /Sera-4.5A-Full-T1This dataset contains 72118 trajectories. Data was generated from the first rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T1.text10K<n<100K2 likes226 downloads7mo agoHugging Face23mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes178 downloads7mo agoHugging Face24jhanglee /youtube-highlights-full YouTube Highlights 完整媒体与标注 本仓库面向数据集协作交付,提供一个可断点续传的完整 tar 文件。解压后即可得到视频、 官方标签转换结果、字段说明和本地可视化检查页。 数据概况 6 个类别:dog、gymnastics、parkour、skating、skiing、surfing 417 个通过 ffprobe 完整性检查的 MP4 315 个 human_mturk 视频:具有 MTurk 人工软投票分数 102 个 weak_match 视频:只有官方自动匹配弱标签 官方清单中另有 1 个当前不可下载的视频,未进入训练标注 19 个已下载视频存在媒体帧数与官方标注帧号差异,保留在数据集中并单独列入复核清单 Linux 下载与解压 BASE_URL="https://huggingface.co/datasets/jhanglee/youtube-highlights-full/resolve/main" wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jhanglee/youtube-highlights-full.tabularn<1K0 likes168 downloads21d agoHugging Face25allenai /Sera-4.5A-Full-T2This dataset contains 66337 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Sera-4.5-Lite-T2 is a subset of this dataset and was used to train SERA-32B-GA. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T2.text10K<n<100K3 likes167 downloads7mo agoHugging Face26Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes147 downloads1y agoHugging Face27greenstainedglass /amc12-full AMC12 Dataset (Research-Oriented) A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks. This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research. 📘 Introduction The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/greenstainedglass/amc12-full.text1K<n<10K0 likes147 downloads3mo agoHugging Face28azaad /longbench-qkv-qwen3-fulltextn<1K0 likes144 downloads5mo agoHugging Face29wanyu /IteraTeR_full_sentPaper: Understanding Iterative Revision from Human-Written Text Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang Github repo: https://github.com/vipulraheja/IteraTeR text100K<n<1M1 likes141 downloads4y agoHugging Face30Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes137 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.