datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Fixer-Train-Editing-CoT-70KFiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.ramanv-image-editing
ramanv-image-editing
Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models.
Size
592,141 total editing pairs
Sources: ultraedit
Schema
Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available).
Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256.
Licenses
MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0
UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.handy-dictation-editing
Handy dictation-editing corpus
Turns a raw dictated transcript into the text the speaker meant to write.
in : um so the meeting is uh moved to friday no wait thursday at three
out: The meeting is Thursday at three.
Three jobs at once, because they are not separable in speech: drop filler words,
repair punctuation and capitalisation, and — the hard one — when the speaker
changes their mind mid-sentence, delete the wording they abandoned and keep only
what they settled on.
Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.editing_promptsFiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/CiaranCw/FiVE-Fine-Grained-Video-Editing-Benchmark.document-editingThis was meant to be training data to teach an LLM to do some basic document editing tasks.
File: wikipedia_word_sub.json
Input: 150 Wikipedia articles + A request to substitute one word for another (usually a synonym)
Output: The same article, with the word substituted as requested
Format: Fastchat
File: wikipedia_err_correct.json
Input: 224 Wikipedia articles with typos and other errors introduced randomly using the python typo library + A request to fix errors
Output:… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/document-editing.adaption-bsad-synbio-gene-editing
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synbio_gene_editing
This dataset contains synthetic entries (IDs 511–530) simulating records for synthetic biology and gene editing research studies. Each entry includes structured fields such as organism type, biosafety level (BSL-1), risk assessment, and global regulatory context formatted for CSV output. The data is designed to represent low-risk genetic engineering projects under WHO… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synbio-gene-editing.Rethinking_Benchmarking_model_editing
