datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineVisionConcatShuffleIFXEmbodied-R1.5-SFT-Dataset
Embodied-R1.5-SFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.LeWM-Datasets
LeWM Datasets
Canonical JPEG95 Lance datasets used by the LeWM-JAX experiments.
Source snapshot: Server 23, /data/dzb/stablewm-data/datasets/.
cube_single_expert.lance
pusht_expert_train.lance
reacher.lance
tworoom.lance
Use these four directories as the shared source of truth when synchronizing
training data to other experiment servers.
ResBIM-IFC
ResBIM-IFC
ResBIM-IFC is a translated version of the ResBIM dataset in IFC format.
The original ResBIM dataset was only available with Revit project files. This repository provides IFC exports or translations of that material so the data can be used in workflows and tools that rely on the open IFC standard.
Contents
The dataset files are stored under data/.
Provenance
Source dataset: ResBIM
Original repository:… See the full description on the dataset page: https://huggingface.co/datasets/tsesterh/ResBIM-IFC.Embodied-R1.5-RFT-Dataset
Embodied-R1.5-RFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.ift-eval-us-real
Eval IFT SD 1.5 — foto asli, US, gender
Gambar hasil generate dan label Gemini untuk arm foto asli (fassabilf/sd15-ift-us-real,
dataset train fassabilf/ift-train-us-real). Bukan run sintetis — itu ada di
fassabilf/ift-test-us-sd15.
split
checkpoint
protokol
n per okupasi
test_ep*
ep5, 10, 15, 20, 25, 30
pilot_sd15 apa adanya: varian framing, --seed-mode paired --seed-base 0, --max-attempts 3, batch 8
100
val_ep*
ep2..ep30 tiap 2 epoch
sama, tapi --max-attempts 1… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/ift-eval-us-real.COSMOSvsi-benchift-eval-us-dpo
Gambar eval — arm DPO (foto asli & sintetis) dan SFT even, US, gender, SD 1.5
Semua gambar di sini dirender model (bukan foto asli), plus label judge
gemini-3.1-pro-preview per gambar. Protokol pilot_sd15: 14 okupasi, n=100/okupasi
(test), 30 prompt/okupasi (val).
folder
model
checkpoint
dpo_sd15/test_ep*
fassabilf/sd15-dpo-us-sd15
ep5/10/15/20
dpo_real/test_ep*
fassabilf/sd15-dpo-us-real
ep5/10/15/20
dpo_real/val_ep*
fassabilf/sd15-dpo-us-real
ep2..20 tiap 2… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/ift-eval-us-dpo.PointBenchiFLYTABiFLYTAB dataset, taken from this link via this comment
I am NOT the original author, and I am not responsible for any content and quality of this dataset!
As far as I am aware, the iFLYTAB dataset is originally introduced in the SEMv2: Table Separation Line Detection Based on Instance Segmentation paper, so all credit goes to original authors (Zhenrong Zhang and Pengfei Hu and Jiefeng Ma and Jun Du and Jianshu Zhang and Huihui Zhu and Baocai Yin and Bing Yin and Cong Liu). I would like to… See the full description on the dataset page: https://huggingface.co/datasets/gvl610/iFLYTAB.Part-Affordance-2Kifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.ego-planVABench-PRoboVQAvabench-vdata_girlslike_iu
data_girlslike_iu
用于图像生成模型训练的人物图像数据集,仅用于个人本地测试与验证。
数据内容
图像数量:36 张
图像格式:PNG
图像质量:整体较好
标注数量:36 个同名 TXT 文件
标注语言:中文
文件命名:iu_XXX.png 与对应的 iu_XXX.txt
每张图片均配有同名文本标注,可用于支持图像与文本配对数据的训练工具。
推荐用途
Z-Image/Krea2 的 LoRA/LOKR 人物或风格训练
图像生成模型的个人研究与实验
图像文本配对数据处理
使用方法
下载后请保持 PNG 图片和同名 TXT 标注文件处于同一目录,再按所使用训练工具的说明配置训练集路径。
valid目录下为验证数据集,可用于onetrainer或ai-toolkit等工具的验证loss。
许可
本数据集以 MIT License… See the full description on the dataset page: https://huggingface.co/datasets/ifmylove2011/data_girlslike_iu.PIO-Benchsharerobot_trajectoryopen-eqaRoborefitmy_data_handsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fr3",
"total_episodes": 8,
"total_frames": 2730,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_hands.pixmo-points-evalHungarianDocQA_IT_IfEvalQA_v2HungarianDocQA_IT_IFEvalQAdata_girlslike_lt
data_girlslike_lt
用于图像生成模型训练的中文人物图像数据集,仅用于个人本地测试与验证。
数据内容
图像数量:32 张
图像格式:PNG
图像质量:整体中等,少数较差
标注数量:32 个同名 TXT 文件
标注语言:中文
文件命名:lt_XXX.png 与对应的 lt_XXX.txt
每张图片均配有同名文本标注,可用于支持图像与文本配对数据的训练工具。
推荐用途
Z-Image/Krea2 的 LoRA/LOKR 人物或风格训练
图像生成模型的个人研究与实验
图像文本配对数据处理
使用方法
下载后请保持 PNG 图片和同名 TXT 标注文件处于同一目录,再按所使用训练工具的说明配置训练集路径。
valid目录下为验证数据集,可用于onetrainer或ai-toolkit等工具的验证loss。
许可
本数据集以 MIT License… See the full description on the dataset page: https://huggingface.co/datasets/ifmylove2011/data_girlslike_lt.xarm6_act_pick2
