datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EditVerseBench
EditVerse
This repository contains the instruction-based video editing evaluation benchmark for EditVerseBench in paper "EditVerse: A Unified Framework for Editing and Generation via In-Context Learning".
Xuan Ju12, Tianyu Wang1, Yuqian Zhou1, He Zhang1, Qing Liu1, Nanxuan Zhao1, Zhifei Zhang1, Yijun Li1, Yuanhao Cai3, Shaoteng Liu1, Daniil Pakhomov1, Zhe Lin1, Soo Ye Kim1*, Qiang Xu2*
1Adobe Research 2The Chinese University of Hong Kong 3Johns Hopkins University *Corresponding… See the full description on the dataset page: https://huggingface.co/datasets/sooyek/EditVerseBench.au-editingImage-Gen-or-Image-Editing
Image Gen or Image Editing
This dataset is designed for text classification of prompts provided by users. It determines whether a prompt is intended for image generation or image editing.
marbella-price-data
Marbella price data
Open datasets from Marbella Wire, an independent price tracker
for Marbella, Spain. Every figure here is published on marbellawire.com first; this
dataset is the machine-readable copy, regenerated from the same data files that render
the pages. It mirrors the canonical GitHub repository
shiftdylson1/marbella-price-data.
Config
What it is
Rows
Canonical page
sunbed-index
High-season price of two sunbeds plus minimum spend at Marbella beach venues… See the full description on the dataset page: https://huggingface.co/datasets/editorwire11/marbella-price-data.flow-edit-cube-triple-tkgrid-seeds30003-40004
Edit-placement (t x K x eb) campaign — OGBench cube-triple, task 2 — seeds 30003 & 40004
Seed scope: this repository contains only seeds 30003 and 40004. It is not the
full seed set for this campaign — seeds 10001 and 20002 were trained on separate hardware
and are not included here. Any per-cell mean computed from this repo alone is an n=2
estimate; see Caveats.
290 training runs from the uedit_place agent: a grid over where in the flow a
value-driven edit is applied (t), how… See the full description on the dataset page: https://huggingface.co/datasets/jaehyeokdoo2/flow-edit-cube-triple-tkgrid-seeds30003-40004.gpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.guide-to-level-measurement-2021-edition-en-77708_combined_synthetic_datadisaster_editedannotations_creators:
expert-generated
language_creators:
found
languages:
en
licenses:
mit
multilinguality:
monolingual
paperswithcode_id: acronym-identification
pretty_name: disaster
size_categories:
10K<n<100K
source_datasets:
original
task_categories:
token-classification
task_ids: []
gene_editing
Gene Editing Dataset
This dataset is part of the Deep Principle Bench collection.
Files
gene_editing.csv: Main dataset file
Usage
import pandas as pd
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("yhqu/gene_editing")
# Or load directly as pandas DataFrame
df = pd.read_csv("hf://datasets/yhqu/gene_editing/gene_editing.csv")
Citation
Please cite this work if you use this dataset in your research.
who-covid-19-epidemiological-update-edition-163
World Health Organization (WHO) Epidemiological Update - Edition 163 (for embeddings)
train.pnf is taken from the WHO website
test.csv was generated by GPT-3.5-turbo
All text is chunked to a length of 500 tokens with 10% overlap.
arXivEdits_edits
Dataset Card for ArXivEdits (Edits)
ArXivEdits is a dataset comprising 751 English scientific papers from arXiv, each with sentence alignments across multiple revisions.
It also includes fine-grained, span-level edits which are annotated with the revision type and the underlying intention for 1000 sentences.
This dataset consists of only the edits subset of the whole dataset. The sentence-aligned papers can be found in this dataset.
Dataset Sources
Check the… See the full description on the dataset page: https://huggingface.co/datasets/miwytt/arXivEdits_edits.pintora-edit-instructdroidnexus-arabic-english-editorial-retrieval-mini
DroidNexus Arabic-English Editorial Retrieval Mini
A compact public query-to-target set built from live DroidNexus coverage to test bilingual editorial retrieval, cross-language discovery, and workflow-aware search ranking.
Why this exists
This dataset is the public artifact layer for DroidNexus Labs. It turns the site's bilingual retrieval analysis into a compact benchmark for editorial search, cross-language discovery, and workflow-aware ranking experiments.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-english-editorial-retrieval-mini.image-editing-model-notes
Image Editing Model Notes
Working notes on image-generation and prompt-based editing models.
I’m mainly interested in what happens after the first good-looking image:
whether the model follows small editing instructions
whether faces and expressions remain consistent
whether untouched objects quietly change
how well models handle text replacement
whether exact object counts are respected
how lighting edits affect skin and image texture
A visually strong result is not always a… See the full description on the dataset page: https://huggingface.co/datasets/drifterAI3000/image-editing-model-notes.StepBackSearch-ds-phi-editiondroidnexus-arabic-editorial-speech-scorecard-mini
DroidNexus Arabic Editorial Speech Scorecard Mini
A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable.
Why this exists
This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.archive-govt-nz-treasury-csv
Archive Govt NZ — Treasury CSV derivative
Simple Viewer-compatible CSV representation of 54 normalized Treasury dataset
metadata records. The Parquet derivative and preservation source archive remain
available separately.
Rich-Txt-Edit-CoVR
Model and Data for CVPR 2025 CoVR Challenge Submission
This repository contains the checkpoints and enriched dataset for our solution to the CVPR 2025 Composed Video Retrieval Challenge.
File Descriptions
stage2.ckpt: Final Model Checkpoint. This is the definitive model used for our final submission. It was fine-tuned with Focal Loss for top-rank performance. Use this file to reproduce our results.
stage1.ckpt: Intermediate Model Checkpoint. This model is the result of… See the full description on the dataset page: https://huggingface.co/datasets/kyrielw/Rich-Txt-Edit-CoVR.SG-SS-dataset-editadoimage_editWELFake_Dataset_Editedorca_editededited_setno_robot_editedAI_Guarded_Code_Editor_Training_Datasetno_robot_editedtest_new_editedQtest-csv-editiontest-csv-editjp_editorials_about_russia_2024
