datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emit-test-dataset
Dataset Card for EMIT-MSeg Dataset
If you use this dataset, please cite our article:
@misc{herec2026fastmethanedetectionpipeline,
title={A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet},
author={Jonáš Herec and Vít Růžička and Rado Pitoňák and Jan Sedmidubsky},
year={2026},
eprint={2606.03675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.03675},
}… See the full description on the dataset page: https://huggingface.co/datasets/onboard-coop/emit-test-dataset.scarlet-test-dataLora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.isu-challenge-dataset
Dataset Card for ISU Challenge Dataset
Dataset Summary
ISU Challenge Dataset is a synthetic, multi-modal in-cabin automotive dataset.
The dataset contains 1000 synchronized samples with:
RGB render
depth (EXR and PNG)
instance segmentation
canny edge map
structured scenario labels
Each sample is linked through a manifest entry and shares the same sample index and name across modalities.
Supported Tasks
This dataset can support:
Semantic… See the full description on the dataset page: https://huggingface.co/datasets/ISU-Test/isu-challenge-dataset.DAEFR_test_datasetsWe evaluate DAEFR on one synthetic dataset CelebA-Test, and two real-world datasets LFW-Test and WIDER-Test.
Datasets
Filename
Short Description
Source
CelebA-Test (HQ)
celeba_512_validation.zip
3000 (HQ) ground truth images for evaluation
RestoreFormer
CelebA-Test (LQ)
self_celeba_512_v2.zip
3000 (LQ) synthetic images for testing
Ourselves
LFW-Test (LQ)
lfw_cropped_faces.zip
1711 real-world images for testing
VQFR… See the full description on the dataset page: https://huggingface.co/datasets/LIAGM/DAEFR_test_datasets.testdataimage-matching-test-datasettest-data
Data for Nunchaku Tests
external_data_test_exampletestdataCoursera_homework_dataset_test
Dataset Card for Homework Test Set for Coursera MOOC - Hands Data Centric Visual AI
This dataset is the test dataset for the homework in the Hands-on Data Centric Visual AI Coursera course.
This is a FiftyOne dataset with 4572 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples'… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Coursera_homework_dataset_test.Coursera_lecture_dataset_test
Dataset Card for Lecture Test Set for Coursera MOOC - Hands Data Centric Visual AI
This dataset is the test dataset for the in-class lectures of the Hands-on Data Centric Visual AI Coursera course.
This is a FiftyOne dataset with 4159 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Coursera_lecture_dataset_test.AdaGaR_testdatatable_rec_test_dataset
表格识别测试集
数据集简介
该数据集包含百度生成工具 20 张有线 20 张无线,wtw 数据集 15, pubnet val 集 20 张,自我零散标注 18 张,共计 93 张表格图片,涵盖多种场景、不同光照条件、不同的图像分辨率。
该数据集可以结合 表格指标评测库-TableRecognitionMetric 使用,快速评测各种表格还原算法。
关于该数据集,欢迎小伙伴贡献更多数据呦!有任何想法,可以前往 issue讨论。
如果遇到标注有误的,还请指出。
数据集支持的任务
可用于自定义数据集下的模型验证和性能评估等。
数据集的格式和结构
数据格式
数据集只有测试集,仅用于客观评估算法表现。
data
└── test
├── images
│ ├── 000cce9ca593055d4618466e823e6d7c.jpg
│ ├── 0aNtiNtRRLqEZ9y6PuShtAAAACMAAQED.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SWHL/table_rec_test_dataset.test-big-dataset
Dataset Card for Danish WIT
Dataset Summary
Google presented the Wikipedia Image Text (WIT) dataset in July
2021, a dataset which contains
scraped images from Wikipedia along with their descriptions. WikiMedia released
WIT-Base in September
2021,
being a modified version of WIT where they have removed the images with empty
"reference descriptions", as well as removing images where a person's face covers more
than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.smart-product-pricing-2025_test_datassynth_data-test
S-SYNTH
S-SYNTH is an open-source, flexible skin simulation framework to rapidly generate synthetic skin models and images using digital rendering of an anatomically inspired multi-layer, multi-component skin and growing lesion model. It allows for generation of highly-detailed 3D skin models and digitally rendered synthetic images of diverse human skin tones, with full control of underlying parameters and the image formation process.
Framework Details
S-SYNTH… See the full description on the dataset page: https://huggingface.co/datasets/didsr/ssynth_data-test.test_datasettext_det_test_dataset
文本检测测试集
数据集简介
该测试集包括卡证类、文档类和自然场景三大类。其中卡证类有82张,文档类有75张,自然场景类有55张。
该数据集可以结合文本检测指标评测库-TextDetMetric使用,快速评测各种文本检测算法。
关于该数据集,欢迎小伙伴贡献更多数据呦!有任何想法,可以前往issue讨论。
数据集支持的任务
可用于自定义数据集下的模型验证和性能评估等。
数据集加载方式
from datasets import load_dataset
dataset = load_dataset("SWHL/text_det_test_dataset")
test_data = dataset['test']
print(test_data)
数据集生成的相关信息
原始数据
数据来源于网络,如侵删。
数据集标注… See the full description on the dataset page: https://huggingface.co/datasets/SWHL/text_det_test_dataset.test-dataset-9
Breakpoint Grounding 55M
Quick start
from datasets import load_dataset
ds = load_dataset("BreakpointAI/breakpoint-grounding-55m", split="train")
ds[0] # {'image': <PIL.Image>, 'image_caption': ..., 'object_captions': [...], 'normalized_boxes': [...], 'img_size_wh': [...]}
Dataset summary
Breakpoint Grounding 55M is, to our knowledge, the largest instance-grounded image–text dataset
released publicly. Every image comes with an image-level caption… See the full description on the dataset page: https://huggingface.co/datasets/BreakpointAI/test-dataset-9.test-dataset-8
Breakpoint Grounding 55M
Quick start
from datasets import load_dataset
ds = load_dataset("BreakpointAI/breakpoint-grounding-55m", split="train")
ds[0] # {'image': <PIL.Image>, 'image_caption': ..., 'object_captions': [...], 'normalized_boxes': [...], 'img_size_wh': [...]}
Dataset summary
Breakpoint Grounding 55M is, to our knowledge, the largest instance-grounded image–text dataset
released publicly. Every image comes with an image-level caption… See the full description on the dataset page: https://huggingface.co/datasets/BreakpointAI/test-dataset-8.test_datasetibbi_test_data
Dataset Card for IBBI Bark Beetle Testing Dataset
Dataset Summary
This dataset is the primary testing and benchmarking set for the ibbi Python package. It contains images of bark and ambrosia beetles used to evaluate the performance of object detection and classification models.
Note: While this dataset serves as the testing set for the ibbi package's evaluation functions, it is hosted on the Hugging Face Hub as the train split. You can access it using… See the full description on the dataset page: https://huggingface.co/datasets/IBBI-bio/ibbi_test_data.DDR_dataset_train_test
🐨 DR Classification Fundus Dataset
This dataset contains retinal fundus images labeled for Diabetic Retinopathy (DR) classification, intended for use in machine learning tasks such as image classification and medical diagnosis support.
📁 Dataset Structure
The dataset is stored under the splits/ directory and is divided into:
train and test folders with train and test.csv correspondingly
vla-data-aug-veo31-test
VLA Data Augmentation (Veo 3.1)
Synthetic video samples for vision-language-action (VLA) data augmentation: 8s clips generated with Veo 3.1 on Vertex AI, with per-clip frames and prompt metadata.
Dataset summary
Model: veo-3.1-generate-001
Duration: 8 seconds per clip, 16:9, no audio
Content: Multi-scene prompts (meetings, care, logistics, office, etc.)
First batch (runs/): 11 prompts × 6 clips each (English prompts); see RUN_PLAN_5D_11PROMPTS_X6.md. 補齊至每題 18… See the full description on the dataset page: https://huggingface.co/datasets/sou35/vla-data-aug-veo31-test.test-datatest_lerobot_dataset
DataArm Dataset
Dataset created with DataArm robot system for LeRobot v3.0 compatibility.
Dataset Description
DataArm teleoperation
Dataset Structure
episodes.json: Episode information
train/: Training data in Parquet format
meta/: Metadata and dataset card
images/: Camera frames (if available)
Features
observation.state: float32 [8] (observation)
observation.gripper_state: float32 [1] (observation)
observation.images.cam_left: uint8 [480, 640, 3]… See the full description on the dataset page: https://huggingface.co/datasets/lr-2002/test_lerobot_dataset.external_data_test_example_v2test_datazalo-ai-2025-public-test-data-v2
Zalo AI Challenge 2025 - RoadBuddy Public Test Data V2set
This dataset contains public test data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-public-test-data-v2.
