datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimScale
Haochen Tian,
Tianyu Li,
Haochen Liu,
Jiazhi Yang,
Yihang Qiu,
Guang Li,
Junli Wang,
Yinfeng Gao,
Zhang Zhang,
Liang Wang,
Hangjun Ye,
Tieniu Tan,
Long Chen,
Hongyang Li
📧 Primary Contact: Haochen Tian (tianhaochen2023@ia.ac.cn)
📜 Materials: 🌐 𝕏 | 📰 Media| 🗂️ Slides | 🎬 Talk (in Chinese)
🖊️ Joint effort by CASIA, OpenDriveLab at HKU, and Xiaomi EV.
🔥 Highlights
🏗️ A scalable simulation pipeline that synthesizes diverse and… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/SimScale.YouTube-Cantonese
Cantonese Audio Dataset from YouTube
This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.Thalia
Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring
Paper | GitHub | Interactive Demo (Colab)
Thalia is a global, multi-modal dataset for volcanic activity monitoring through Satellite-based Interferometric Synthetic Aperture Radar (InSAR) imagery. Building upon the Hephaestus dataset, Thalia provides higher-resolution, multi-source, and multi-temporal data in a machine-learning-ready format.
Dataset Overview
Thalia consists of 38 spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Thalia.OriAnyV2_Train_Render
Orient Anything V2 Dataset
Project Page | Paper | GitHub
Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This repository contains the training data (final rendering data) used for the model.
Sample Usage
Below is a snippet to run inference using the model and data logic, as found in the official GitHub repository:
import numpy as np
from PIL importImage
import torch
import… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/OriAnyV2_Train_Render.OriPID
Summary
This is the dataset proposed in our paper Origin Identification for Text-Guided Image-to-Image Diffusion Models (ICML 2025).
Download
Training
You can download the images:
wget https://huggingface.co/datasets/WenhaoWang/OriPID/resolve/main/training/sd2_d_multi.tar.part_0{0..9}
cat sd2_d_multi.tar.part_* > sd2_d_multi.tar
tar -xvf sd2_d_multi.tar
Or you can directly download the features extracted by VAE in Stable Diffusion 2:
wget… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/OriPID.Gen-nuScenesYouTube-English
English Audio Dataset from YouTube
This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.bmc_original_1Mwds_flickr30k_orderwds_vtab-dsprites_label_orientationwds_coco_orderCombineCorpus_6_Orgwds_dsprites_label_orientationoro_depth_rewardwds_vtab-dsprites_label_orientation_test
dSprites Orientation (Test set only)
Original paper: beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Homepage: https://github.com/deepmind/dsprites-dataset
Bibtex:
@misc{dsprites17,
author = {Loic Matthey and Irina Higgins and Demis Hassabis and Alexander Lerchner},
title = {dSprites: Disentanglement testing Sprites dataset},
howpublished= {https://github.com/deepmind/dsprites-dataset/},
year = "2017",
}
ori_coco_evalbridge_orig_text_20kscenaverseEuroSAT_RGB_Ctrain_original_sd3.5-lzunorigin_data_20260221_raw_img
