datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.stackv2_edu_filtered
Stack V2 Edu
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
editorai-telemetryhplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.Inter-Edit-Train
Inter-Edit-Train
Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing.
This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance.
Highlights
1,099,964 image editing pairs
610,186 unique source images
Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.IndustryCorpus_education[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_education.llm_pt_leaderboard_resultsfineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.STER
STER: Zero-shot 3D Geometric Entity Resolution Benchmark
Multi-city, cross-LoD 3D building matching benchmark for the NS-D2S paper (AAAI 2026).
Strictly follows the 3dSAGER (SIGMOD 2026) methodology and data format.
Dataset Overview
Dataset
City
Country
Buildings
LOD Source
Urban Typology
amsterdam
Amsterdam
NL
123,259
3DBAG LOD1.2/1.3/2.2
Historic canal city
rotterdam
Rotterdam
NL
152,694
3DBAG LOD1.2/1.3/2.2
Post-war modern
hague
Den Haag
NL
181… See the full description on the dataset page: https://huggingface.co/datasets/eduzrh/STER.TransWeb-Edu-SpanishTransWeb-Edu-FrenchJQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.hpltv2-llama33-edu-annotation
HPLT version 2.0 educational annotations
This dataset contains annotations derived from HPLT v2 cleaned samples.
There are 500,000 annotations for each language if the source contains at least 500,000 samples.
We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier.
Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.TransWeb-Edu-GermanDeepSWEGym2-Edu
Dataset Description
This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 291.74kb, a total uncompressed size of 14.93GB, and a total of 53649 examples.
Dataset Details
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Edu.TransWeb-Edu-Englishamc12-full
AMC12 Dataset (Research-Oriented)
A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks.
This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research.
📘 Introduction
The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/edev2000/amc12-full.commit-msg-edits
✍️ Commit Message Edits Dataset
This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS.
You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.DIM-Edit
[ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
Ziyun Zeng, David Junhao Zhang, Wei Li,
and Mike Zheng Shou
📰 News
[2026-05-12] The DIM project page is available.
[2026-01-26] 🎉 DIM is accepted to ICLR 2026!
[2025-10-08] 🚀 Released the DIM-Edit dataset and the DIM-4.6B-T2I/ DIM-4.6B-Edit models.
[2025-09-02] 📝 The DIM paper is released on arXiv.
🌟 Highlights
🧠 Rebalanced… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/DIM-Edit.fineweb-eduEduQS
EduQS Dataset
EduQS is a multi-subject, multi-grade-level dataset for Chinese visual question answering in the K12 education domain. It contains high-quality structured question data with accompanying illustrative images and answer keys.
💡 Highlights
Covers subjects: Biology, Chemistry, Physics, History, Geography, Math
Grade levels: Middle School and High School
Question types: fill-in-the-blank, multiple-choice, open-ended
Includes annotated solutions, side… See the full description on the dataset page: https://huggingface.co/datasets/chaosY/EduQS.fineweb_edu_10B_for_crypto-LLMabeja-cc-ja-edu-10percentnornikel-metallurgy-vl-dataset
Nornikel Metallurgy VL Dataset (SFT / DPO / GRPO)
Датасет для дообучения мультимодальной модели Qwen3-VL по схеме
SFT → DPO → GRPO в предметной области металлургии, горного дела и
обогащения полезных ископаемых. Построен из корпуса технических документов
(PDF-книги/сборники, DOCX-отчёты, PPTX-презентации, XLSX-таблицы) и
изображений (схемы, диаграммы, таблицы).
Конфигурации (config_name)
config
train
validation
назначение
sft
111 351
12 372… See the full description on the dataset page: https://huggingface.co/datasets/brics-edtech/nornikel-metallurgy-vl-dataset.DeepSWEGym-Edu
Dataset Description
This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Edu.
