datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.watercolour-reference-pool
Watercolour reference pool
The reference paintings that define the reward in the watercolour RL environment: an
agent writes a p5.brush sketch, the sketch is
rendered, and a vision judge compares the render against paintings sampled from this pool.
What the pool contains is the reward function. Replace it and you have changed what
the environment rewards, without touching a line of code.
178 paintings in two tiers, each with the JavaScript source that produced it.
tier… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-reference-pool.asset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.ssim-reference-videospeg_rand_05_01_cam_reference_cam0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_cam_reference_cam0.peg_rand_05_01_tip_reference_cam0_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam0_1.peg_rand_05_01_tip_reference_cam0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam0.peg_rand_05_01_tip_reference_cam1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam1.phisat2-ortho-referenceNegima-manga-reference-chapters
Negima Manga Reference Chapters (Public)
Personal reference dataset for AI image & video generation (Artlist Seedance 2.0, Kling, LoRAs, etc.).
Mahou Sensei Negima! (Negima!) by Ken Akamatsu
UQ Holder! (sequel series) — Chapters added
Perfect for consistent characters (Asuna, Setsuna, Konoka, Touta, Kirie, etc.), magic circles, pactio cards, Ensis Exsequens slashes, wind blades, immortal fights, and Ken Akamatsu art style.
english versions coming soon
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/Zentoria/Negima-manga-reference-chapters.peg_rand_05_01_cam_reference_cam1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_cam_reference_cam1.style-reference-captionsdog-shelter-referenceterrain-referenced-3d-glacier-mapping-product
Terrain-referenced Glacier Mapping Product
This repository provides the terrain-referenced glacier-area mapping product generated for the manuscript. The product is openly available through Hugging Face with DOI: 10.57967/hf/9900.
The archive contains regional mapping outputs, oblique terrain-visualization products, metadata files, and tabular glacier-area summaries. Glacier masks generated by Prithvi-SDT are linked with Copernicus DEM terrain information and RGI 7.0 glacier… See the full description on the dataset page: https://huggingface.co/datasets/yyhw/terrain-referenced-3d-glacier-mapping-product.visual_instruction_tuning_ID_referencepyOpenFOAM-reference-data
pyOpenFOAM Reference Data & Validation Results
OpenFOAM-13 reference simulation data and pyOpenFOAM validation results for pyOpenFOAM — a pure Python/PyTorch reimplementation of OpenFOAM with GPU acceleration and automatic differentiation.
Dataset Summary / 数据摘要
Property
Value
Total reference cases
257
Validated cases
233 (90.7%)
Categories
21
Source
OpenFOAM v11/v13
Reference data size
2.42 GB
pyOpenFOAM results
3.3 MB
Field files analyzed… See the full description on the dataset page: https://huggingface.co/datasets/AlanZee/pyOpenFOAM-reference-data.Anime_Character_Transfer_and_Reference_Dataset
Anime Character Transfer and Reference Dataset
This dataset is designed for anime-style character transfer, reference-based image editing, and multi-reference character consistency experiments.
Each sample contains a source/reference pair and a text prompt. Most samples also include a generated target image and a metadata file. A small number of samples are kept as reference-only entries, so they can still be used for reference-pair tasks or future target completion.… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/Anime_Character_Transfer_and_Reference_Dataset.Design2Code_human_eval_reference_vs_gpt4vFind more details in our paper.
reference_datasetsenryu-test-with-references
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/senryu-test", split="test")
概要
川柳投稿サイトの『写真川柳』と『川柳投稿まるせん』のクロールデータです。
以下のページからクロールし、原本のHTMLファイルと構造化処理を行った結果を格納しました。
https://www.homemate-research.com/senryu/photo/
https://marusenryu.com/
このデータは以下の2タスクが含まれます。
image_to_text: 画像でお題が渡され、それに対する回答を返します。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
それぞれの量は以下の通りです。
タスク
お題数(画像枚数)
回答数
うち委員が用意したお題
image_to_text
70
140
7
text_to_text
30
60… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/senryu-test-with-references.Horror-Reference-Datastyle-reference-captionsvisual-reference-assetsogiri-test-with-references
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/bokete-ogiri-test", split="test")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。
タスク
お題数(画像枚数)
回答数
うち委員が用意したお題… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-test-with-references.lego_annotated_reference_debug_fixed_radius002_20260610
LEGO annotated reference debug fixed radius 0.02
Visual debug export for the custom_3cam annotated-reference ablation after increasing inner dot radius to 0.02 * min(H, W), which is 3 px at 160p and matches the older ajaysri/lego_stack_4x4_v1_teleop_frontdot_backwrist_lerobot_v3 LeRobot export dot radius.
Each row shows raw reference RGB, fixed annotated reference without jitter, and fixed annotated reference with 25% bbox-size dot jitter. Dots use a white outline.
pyOpenFOAM-reference-data
pyOpenFOAM Reference Data & Validation Results
OpenFOAM-13 reference simulation data and pyOpenFOAM validation results for pyOpenFOAM — a pure Python/PyTorch reimplementation of OpenFOAM with GPU acceleration and automatic differentiation.
Dataset Summary / 数据摘要
Property
Value
Total reference cases
257
Validated cases
233 (90.7%)
Categories
21
Source
OpenFOAM v11/v13
Reference data size
2.42 GB
pyOpenFOAM results
3.3 MB
Field files analyzed… See the full description on the dataset page: https://huggingface.co/datasets/isunme/pyOpenFOAM-reference-data.peg_rand_05_01_cam_reference_cam0_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_cam_reference_cam0_1.TreeOil_TorquePhysics_CorePrinciples_ReferenceGuide📘 README: Principles of Torque-Based Scientific Art Authentication
Dataset: The Tree Oil Painting – Torque Analysis and X-Ray Forensics
🎯 Purpose
This dataset presents a multi-layered forensic investigation of The Tree Oil Painting, using torque-based AI analysis, X-ray imaging, and scientific pigment studies. The goal is to identify the artist's physical signature through measurable physics, not surface appearance.
The dataset adheres to the following fundamental principles of scientific… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_TorquePhysics_CorePrinciples_ReferenceGuide.Veronica_Reference_PhotoGPT-Multi-Reference-Edit
