datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniEdit-Filtered-1.2M
OmniEdit
In this paper, we present OMNI-EDIT, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) OMNI-EDIT is trained by utilizing the supervision
from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality.
📃Paper | 🌐Website |… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/OmniEdit-Filtered-1.2M.instructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.filtered-wit
Filtered WIT, an Image-Text Dataset.
A reliable Dataset to run Image-Text models.
You can find WIT, Wikipedia Image Text Dataset, here
Data was taken from dalle-mini/wit
Author
Aarush Katta
Data Structure
The data is stored as tars, containing 10,000 samples per tar.
The parquets contain the metadata of each tar, which was crated using this script
Each tar contains a .jpg, .txt, and .json.
The image is stored in .jpg, the caption in .txt. and the metadata in… See the full description on the dataset page: https://huggingface.co/datasets/laion/filtered-wit.laion_synthetic_filtered_large_part3laion_synthetic_filtered_large_part1laion_synthetic_filtered_large_part2product-photography-v1-tiny-prompts-tasks-collage-filteredpuzzle-hle-filteredlaion_synthetic_filtered_large_part4ULVR-filtered
ULVR-filtered
Filtered subset of RuoliuYang/ULVR_v2_clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only input_image, but correctly once the intermediate_image_* were also provided (judged by Qwen3-VL-32B-Instruct). Same schema / subsets / train-split structure as the source.
subset
rows
scene_graph
3522
edge
1394
depth
537
segmentation
1328
bbox_highlight
15186
bbox_crop
15260
text_cot
27158
helper_interleaved… See the full description on the dataset page: https://huggingface.co/datasets/williamium/ULVR-filtered.pdf_images_filtered
Image Dataset with Parquet Format
This dataset contains images with their IDs in parquet format for efficient loading.
Dataset Structure
Each configuration (language) contains:
image: PIL Image object
id: String identifier
Languages
Arabic (ar): 955 images
Bengali (bn): 932 images
German (de): 940 images
English (en): 932 images
Spanish (es): 951 images
French (fr): 947 images
Gujarati (gu): 949 images
Hindi (hi): 891 images
Italian (it): 1,005 images… See the full description on the dataset page: https://huggingface.co/datasets/v1v1d/pdf_images_filtered.ccs_synthetic_filtered_largeDanbooru-2024-Filtered-1Msvg-stack-filtered
Dataset Card for svg-stack-filtered
This is an attempt to replicate the dataset used for SFT in the paper
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Processed:
Optimized with svgo precision=2
Rasterized with cairosvg[^cairo]
[^cairo] cairosvg doesn't implement all svg features, but matches how the original paper
Filtered based on some heuristics:
Removed any svg that couldn't be rendered with cairosvg (~30%)
Removed solid-color images
Removed some… See the full description on the dataset page: https://huggingface.co/datasets/darknoon/svg-stack-filtered.instructpix2pix-clip-filtered-upscaledCONCEPTUAL_CAPTIONS_HU_FILTEREDpixmo-points-filtered_0-20_imgContainedinstructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/guyue-wa/instructpix2pix-clip-filtered.ai2thor-random-views-20k-3obj-filteredamazon-all-beauty-filtered-limitedreal_0_put_bowl_filtered_raw_frames
real_0_put_bowl_filtered
Task: "put the bowl on the plate"
Type: training (filtered)
Robot: Franka FR3
Cameras: observation.images.primary, observation.images.wrist (image, 256x256) @ 15 FPS
Statistics
Metric
Value
Episodes
53
Total frames
16675
Avg frames/episode
314
FPS
15
Format
LeRobot v3.0
Features
Feature
Type
Shape
observation.images.primary
image
[256, 256, 3]
observation.images.wrist
image
[256, 256, 3]… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/real_0_put_bowl_filtered_raw_frames.of_filtered_splitbenchmark_coco_filteredCOCO Benchmark Dataset Description and Metadata
MPM_train_filteredosatlas-fineweb-images-filtered-1datacomp-small-filtered
Dataset Card for "datacomp-small-filtered"
This is the DataComp-small dataset with CLIP-large-patch14 image embeddings added, as well as:
captions filtered for English using a FastText model
captions filtered to have at least complexity of 1
rl_expert_franka_dataset_for_pi0_1M_filtered_for_pickupsai2thor-perspective-qa-800-to-400-human-filtered-v2UNO1m-filtered-splitdatacap-recomp-1M-download-filtered
