datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chitralekha
Chitralekha
Dataset Details
Dataset Version
Some of the fonts do not have proper letters/rendering of different telugu letter combinations. Those have been removed as much as I can find them. If there are any other mistakes that you notice, please raise an issue and I will try my best to look into it
Dataset Description
This extensive dataset, hosted on Huggingface, is a comprehensive resource for Optical Character Recognition (OCR) in the Telugu… See the full description on the dataset page: https://huggingface.co/datasets/gksriharsha/chitralekha.Chicks4FreeID
Dataset Card for Chicks4FreeID
The very first publicly available dataset for chicken re-identification.
1 Dataset Details
1.1 Dataset Description
The Chicks4FreeID dataset contains top-down view images of individually segmented and annotated chickens (with roosters and ducks also possibly present and labeled as such). 11 different coops with 54 individuals were visited for manual data collection. Each of the 677 images depicts at least one chicken. The… See the full description on the dataset page: https://huggingface.co/datasets/dariakern/Chicks4FreeID.chinese-painting-collection
Chinese Painting Collection
91,438 images of Chinese paintings with bilingual (Chinese/English) VLM captions,
calligraphy OCR transcriptions, view-type classification, a long caption on part
of the set, and foreign-object detection with one final crop box per image.
Sources
Component
Source
Images
Image license
npm_tw_c0–npm_tw_c3
National Palace Museum (Taipei) Open Data
86,666
Taiwan Open Government Data License v1 (attribution required)
met_china… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/chinese-painting-collection.LLaVA-Instruct-600K-Chinese
仿照 LLaVA-Instruct-150K ,使用 Qwen2.5-VL-32B-Instruct 合成的用于微调中文VLM的数据;也可以与英文数据集混合使用,训练多语言VLM
任务类型为基于单张图片的问答和对话,每个样本都对应一张不同的图片,其中大部分图片包含中文字符,更适合中文场景下视觉语言模型的训练。
图片从各类中文网站上爬取
包含3类任务:日常对话、复杂推理、描述图片。日常对话通常是5轮对话,其余任务是1轮对话。
每种任务的数量如下:
任务类型
数量
日常对话
247,431
复杂推理
194,646
描述图片
199,791
用于生成对话数据的prompt如下
日常对话
设计一个你和一个询问这张照片的人之间的对话。答案应该是视觉AI助手看到图像并回答问题的语气。
你需要提出不同的问题并给出相应的答案。问题可以包括询问图像视觉内容的问题,包括对象类型、对象计数、对象动作、对象位置、对象之间的相对位置等。必须是有明确答案的问题,即
(1) 人们可以在图像中明确看到问题所问的内容,并且可以自信地回答;
(2)… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/LLaVA-Instruct-600K-Chinese.chinese_landscape_paintings
Dataset Card for "chinese_landscape_paintings"
More Information needed
imagenet-lt-v2reddit
Dataset Card for "reddit"
More Information needed
CHIP
CHIP: A multi-sensor dataset for 6D pose estimation of chairs in industrial settings
🏠 Homepage
📄 Paper
Introduction
Accurate 6D pose estimation of complex objects in 3D environments is essential for effective robotic manipulation. Yet, existing benchmarks fall short in evaluating 6D pose estimation methods under realistic industrial conditions, as most datasets focus on household objects in domestic settings, while the few available industrial… See the full description on the dataset page: https://huggingface.co/datasets/FBK-TeV/CHIP.zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.dancing-chibi-figures
Dancing Chibi Figures — v0.1
One template Q-version (chibi) character, animated by 1,430 motion clips, rendered with exact labels and motion-grounded captions — and paired frame-for-frame with Dancing Stick Figures.
1,423 clips · 6 s @ 20 fps · 128×128 RGBA · 514,800 frames · 143 text prompts × 10 seeds × 3 cameras ·
every frame carries the 3D skeleton, camera, depth, camera-space normals, part segmentation, motion events and five
levels of caption.
One row per motion group;… See the full description on the dataset page: https://huggingface.co/datasets/sprited/dancing-chibi-figures.CHIRLA
Dataset Card for CHIRLA
CHIRLA (Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis) is a long-term, multi-camera person Re-Identification (Re-ID) and tracking dataset. It spans 7 months, 7 cameras, 22 identities, and ~1M identity-annotated bounding boxes across ~596k frames, captured in connected indoor environments.
Dataset Details
Dataset Description
CHIRLA targets long-term appearance change (e.g., clothing changes… See the full description on the dataset page: https://huggingface.co/datasets/bdager/CHIRLA.editscore-rl-train
Introduction
Training data for OmniGen2 Online-RL using EditScore.
Usage
# meta file: rl.jsonl
# images:
cat images_part_* > images.tar.gz && tar -xzvf images.tar.gz
Citation
@article{luo2025editscore,
title={EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling},
author={Xin Luo and Jiahao Wang and Chenyuan Wu and Shitao Xiao and Xiyan Jiang and Defu Lian and Jiajun Zhang and Dong Liu and… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/editscore-rl-train.laion-high-resolution-chinese
laion-high-resolution-chinese
简介 Brief Introduction
取自Laion5B-high-resolution多语言多模态数据集中的中文部分,一共2.66M个图文对。
A subset from Laion5B-high-resolution (a multimodal dataset), around 2.66M image-text pairs (only Chinese).
数据集信息 Dataset Information
大约一共2.66M个中文图文对。大约占用381MB空间(仅仅是url等文本信息,不包含图片)。
Homepage: laion-5b
Huggingface: laion/laion-high-resolution
下载 Download
mkdir release && cd release
for i in {00000..00015}; do wget… See the full description on the dataset page: https://huggingface.co/datasets/wanng/laion-high-resolution-chinese.laion2B-multi-chinese-subset
laion2B-multi-chinese-subset
Github: Fengshenbang-LM
Docs: Fengshenbang-Docs
简介 Brief Introduction
取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。
A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese).
数据集信息 Dataset Information
大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。
Homepage: laion-5b
Huggingface: laion/laion2B-multi
下载 Download
mkdir laion2b_chinese_release && cd laion2b_chinese_release
for i in {00000..00012}; do… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.object_in_bowl_64chinese_fonts_common_128x128
Dataset Card for "chinese_fonts_common_128x128"
More Information needed
Chinese-Multimodal-Instruct
中文(视觉)多模态指令数据集
💻 Github Repo
本项目旨在构建一个高质量、大规模的中文(视觉)多模态指令数据集,目前仍在施工中 🚧💦
[!Important]
本数据集仍处于 WIP (Work in Progress) 状态,目前 Dataset Viewer 展示的是 100 条示例。
初步预计规模大约在 1~2M(不包含其他来源的数据集),均为多轮对话形式。
[!Tip]
[2025/05/05] 图片已经上传完毕,后续文字部分正等待上传。
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
chinese_fonts_common_512x512
Dataset Card for "chinese_fonts_common_512x512"
More Information needed
traditional-chinese-ocr-synthetic
Traditional Chinese OCR Synthetic Dataset
A large-scale synthetic dataset containing 4.1 million image-text pairs specifically designed for Traditional Chinese historical document recognition.
Dataset Overview
Existing large-scale Traditional Chinese OCR datasets (e.g., TCSynth) are primarily designed for scene text recognition, characterized by:
Horizontal layouts
Short text sequences (2-5 characters on average)
Modern commonly-used characters
These characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-ocr-synthetic.Chinese-SimpleVQA
Overview
🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper 中文 | English
Dataset
1.chinese_simple_vqa.jsonl(the image is in url format)
2.chinese_simplevqa.parquet (the image is in base64 format and can be downloaded)
Chinese SimpleVQA is the first factuality-based visual question-answering benchmark in Chinese, aimed at assessing the visual factuality of LVLMs across 8 major topics and 56 subtopics. The key features of this benchmark include a focus… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleVQA.AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.… See the full description on the dataset page: https://huggingface.co/datasets/chintalaswathi/AI-vs-Deepfake-vs-Real-Resized-Aug.anny-dress-on-stage-train
anny-dress-on-stage-train
Dress-on edits of ANNY parametric bodies with second-hand garment photos, edited by
VoxHammer (training-free 3D latent editing on TRELLIS-image-large), rendered by
Mitsuba 3 over a seeded Hammersley camera sequence, and scored with the MaskScore
geometric metric against the source's own decode.
MaskScore-shaped ETNF: dress_on (root), dress_on_candidates (rank1 own garment,
rank3 wrong garment, rank5 source decode = the floor), dress_on_scores (per view… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/anny-dress-on-stage-train.chinese_fontsduplii-synthetic-15kfood_chinese_2017
Dataset Card for "food_chinese_2017"
More Information needed
robocasa_20260430T030150Z_full_run_distribute_chicken ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_distribute_chicken.chimp-detectionthree-mountain-scaling
ThreeMountain_Scaling
Segment
Meaning
GO
Geometric Object — indicates the object type used (e.g., GO for geometric, RO for real objects).
L / Arc
Object Arrangement — defines how objects are arranged spatially. L means L-shape arrangement; Arc means objects are placed in an arc.
RC
Random Character Position — RC = True: character position is randomized.
FC
Fixed Character Position — FC = True: character stays fixed.
RS
Random Scale — RS = True: objects are… See the full description on the dataset page: https://huggingface.co/datasets/grow-ai-like-a-child/three-mountain-scaling.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.
