datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.arvo-vulnsmith-full
ARVO CyberGym-format smoke dataset
This dataset is shaped to be loaded by Harbor's CyberGym adapter.
It contains 10 ARVO tasks that are outside the original CyberGym set.
arxiv-tex-corpus-fullarxiv-tex-corpus-full (80GB)
Large-scale LaTeX corpus from arXiv (math, CS, physics, statistics)
📄 Paper: https://arxiv.org/abs/2602.17288
📚 Overview
arxiv-tex-corpus-full (80GB) is a large-scale dataset of LaTeX source content extracted from papers hosted on arXiv.
This version contains approximately 80GB of structured JSONL data, restricted to the following arXiv categories:
math
cs
hep-th
hep-ph
quant-ph
stat.ML
stat.TH
The dataset is designed for research in:
Large… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-full.amc12-full
AMC12 Dataset (Research-Oriented)
A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks.
This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research.
📘 Introduction
The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/edev2000/amc12-full.yelp_review_fullprostate128_t2_anatomy_nnUNet_3d_fullres_20_epochDeepSWEGym2-Full
Dataset Description
This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples.
Dataset Details
Curated by: MoreThought
Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.prostate158_nnUNet_3d_fullres_20_epochvertebrate-v1-issue473-fullwindow-cds-random-val
marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val
CDS full-window vertebrate projection sequences for the issue #473 random
validation control. The source is the immutable issue #417 accepted-sequence
table.
The split uniformly samples 16,384 original-orientation CDS rows
without replacement using seed 42. Sampling occurs before
reverse-complement augmentation. Selected rows are removed from training;
reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.DeepSWEGym-Full
Dataset Description
This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total.
Dataset Details
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.ruler-full
RULER Benchmark (Full)
Complete RULER benchmark dataset with all 13 tasks across 6 context lengths (4K to 128K tokens).
Overview
Metric
Value
Total Samples
78,000 (39,000 per variant)
Tasks
13
Context Lengths
4K, 8K, 16K, 32K, 64K, 128K
Samples per Config
500
Variants
memwrap, plain
Tasks
Retrieval (NIAH - Needle in a Haystack)
niah_single_1, niah_single_2, niah_single_3 - Single needle variants
niah_multikey_1… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/ruler-full.livesqlbench-base-full-v1
🚀 LiveSQLBench-Base-Full-v1
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks.
🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral)
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-full-v1.cybersecurity_full_question_answersFDAbench-Full
v1.1 Update (2026-08-06) — multiple split
Strengthened the cross-source requirement that multiple-choice tasks are designed
around (selecting all correct options should require integrating both the SQL
result and the retrieved documents): 264 of 760 tasks were revised, with task IDs,
databases, and gold SQL unchanged. Documents-only accuracy drops from 61.5% to
38.7% while full-evidence accuracy stays at 80.6% (3 frontier models, strict
exact set match).
Diversified the number… See the full description on the dataset page: https://huggingface.co/datasets/FDAbench2026/FDAbench-Full.bird-interact-full🌐 Website • 📄 Paper (ICLR 2026 Oral) • 💻 GitHub • 🗄️ bird-interact-lite • 🗄️ bird-interact-full • 🗄️ LiveSQLBench
🧸 Overview
BIRD-INTERACT, an interactive text-to-SQL benchmark, re-imagines Text-to-SQL evaluation via lens of dynamic interactions which is built on top of single-turn unambiguous T2S tasks from LiveSQLBench.
The environment blends a hierarchical knowledge base, database documentation and a function-driven user simulator to recreate authentic enterprise… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-interact-full.vqav2-full-metadatavertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.flan2021-full
Task Name
FLAN-2021 -> 70
{
"ag_news_subset": 108497,
"ai2_arc/ARC-Challenge": 829,
"ai2_arc/ARC-Easy": 1927,
"aeslc": 13187,
"anli/r1": 15361,
"anli/r2": 41133,
"anli/r3": 91048,
"bool_q": 8343,
"cnn_dailymail": 259607,
"coqa": 6456,
"cosmos_qa": 22996,
"definite_pronoun_resolution": 1079,
"drop": 70045,
"fix_punct": 25690,
"gem/common_gen": 60936,
"gem/dart": 56724,
"gem/e2e_nlg": 30337,
"gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows.
Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos.
Rows: 350
Selected repos: 19
Deduped overlap capacity: 468
Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
InternVid-Full
InternVid
InternVid-Full
We present InternVid-230M, a full set of this dataset, consisting of 230 million video clips, with generated high-quality captions for publicly available web videos.
Download
The 230M samples are provided in jsonlines file. Columns include the videoID, timestamps, generated caption and their UMT similarity scores.
How to Use
from datasets import load_dataset
dataset = load_dataset("OpenGVLab/InternVid-Full")
Method… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVid-Full.sets_lego_omr_full
Dataset Card for sets_lego_omr_full
This dataset combines official LEGO sets from LDRAW OMR with metadata from Rebrickable. Each entry contains the full MPD file as a string plus associated metadata such as set number, name, theme, year, and parts count. It is intended for building LLM fine-tuning datasets for LEGO model generation tasks.
Dataset Details
Dataset Sources
OMR files: LDraw OMR Library
Metadata: Rebrickable
MPD file format: LDraw File Format… See the full description on the dataset page: https://huggingface.co/datasets/DylanRiden/sets_lego_omr_full.Sera-4.5A-Full-T1This dataset contains 72118 trajectories. Data was generated from the first rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase from the current trajectory
func_name: Name of function sampled from codebase to start the pipeline
func_path: File path to the sampled function
problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T1.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.youtube-highlights-full
YouTube Highlights 完整媒体与标注
本仓库面向数据集协作交付,提供一个可断点续传的完整 tar 文件。解压后即可得到视频、
官方标签转换结果、字段说明和本地可视化检查页。
数据概况
6 个类别:dog、gymnastics、parkour、skating、skiing、surfing
417 个通过 ffprobe 完整性检查的 MP4
315 个 human_mturk 视频:具有 MTurk 人工软投票分数
102 个 weak_match 视频:只有官方自动匹配弱标签
官方清单中另有 1 个当前不可下载的视频,未进入训练标注
19 个已下载视频存在媒体帧数与官方标注帧号差异,保留在数据集中并单独列入复核清单
Linux 下载与解压
BASE_URL="https://huggingface.co/datasets/jhanglee/youtube-highlights-full/resolve/main"
wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jhanglee/youtube-highlights-full.Sera-4.5A-Full-T2This dataset contains 66337 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Sera-4.5-Lite-T2 is a subset of this dataset and was used to train SERA-32B-GA.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase from the current trajectory
func_name: Name of function sampled from codebase to start the pipeline
func_path:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T2.tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.amc12-full
AMC12 Dataset (Research-Oriented)
A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks.
This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research.
📘 Introduction
The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/greenstainedglass/amc12-full.longbench-qkv-qwen3-fullIteraTeR_full_sentPaper: Understanding Iterative Revision from Human-Written Text
Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang
Github repo: https://github.com/vipulraheja/IteraTeR
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.
