datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.Movie101
Movie101
[!NOTE]
Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2
Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie.
The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.youtube_caption_yue
YouTube ASR Caption Dataset (Cantonese)
This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs.
What’s included
Segments where the ASR output is identical to the original caption — likely clean.
Segments where differences are only homophones (同音字) or English words — likely ASR mistakes.
This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.SlimPajama-6B_km-ip-d512celeba
CelebA dataset
A copy of celeba dataset.
https://mmlab.ie.cuhk.edu.hk/projects/CelebA.html
How to use
Download data
huggingface-cli download --local-dir /path/to/datasets/celeba --repo-type dataset Yuehao/celeba
unzip /path/to/datasets/celeba/img_align_celeba.zip -d /path/to/datasets/celeba
Load data via torchvision.datasets.CelebA
torchvision.datasets.CelebA(root='/path/to/datasets')
common_voice_22_yue2025-08-03 Update:
Use MP3 instead of WAV
All Right reserved by mozilla-foundation
yue_emo_speech
Cantonese Emotional Speech
Crawled from YouTube and RTHK, this dataset contains 1,000 hours of Cantonese speech, each labeled with one of the following emotions: angry, disgusted, fearful, happy, neutral, other, sad, or surprised. The dataset also includes the confidence of the emotion label. The audio files are denoised with resemble-enhance. The transcriptions are generated by SenseVoiceSmall, and deduplicated using MinHash.
speechio_test
SpeechIO ASR Test Sets (parquet)
Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark,
re-exported from yuekai/speechio (Lhotse cuts) into standard
HuggingFace parquet with embedded 16 kHz audio.
27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split.
~43k utterances, ~66 hours total, evaluation only.
Columns
column
type
note
segment_id
string
utterance id
speaker
string
speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.InstructS2S-200Kyue-logiqa
Dataset Card for Cantonese LogiQA
This dataset is a Cantonese translation of jiacheng-ye/logiqa-zh. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"context": "有啲廣東人唔鍾意食辣椒。所以,有啲南方人唔鍾意食辣椒",
"query":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-logiqa.forecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models"
This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.multi_en_qwen3_omni_sglang_regeneratedMMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.GuardReasonerTrain
GuardReasonerTrain
GuardReasonerTrain is the training data for R-SFT of GuardReasoner, as described in the paper GuardReasoner: Towards Reasoning-based LLM Safeguards.
Code: https://github.com/yueliu1999/GuardReasoner/
Usage
from datasets import load_dataset
# Login using e.g. `huggingface-cli login` to access this dataset
ds = load_dataset("yueliu1999/GuardReasonerTrain")
Citation
If you use this dataset, please cite our paper.
@article{GuardReasoner… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasonerTrain.YuE2-MusicYue-Benchmark
How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models
Homepage: https://github.com/jiangjyjy/Yue-Benchmark
Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark
Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models.
Introduction
The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.CV3-Evalcommon_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps
TODO: convert to MP3, reduce size
AlphaPanda_checkpointsforecastingDataset from "Approaching Human-Level Forecasting with Language Models"
This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.seed_tts_cosy2aops_part1Multimodal-Yue-Benchmark
Multimodal Yue Benchmark
Cantonese audio + text benchmark derived from BillBao/Yue-Benchmark (Yue-GSM8K & Yue-MMLU-style tasks). We kept the same task content in Cantonese and added TTS for three Cantonese speakers (hiugaai, hiumaan, wanlung).
Subsets & splits
Config name
Task
Speaker
Splits
mmlu_*
multiple-choice (MMLU-style)
per speaker
train, test
gsm8k_*
math word problems (GSM8K-style)
per speaker
train, test
Example:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/J017athan/Multimodal-Yue-Benchmark.alimeeting_aishell4_training_whisper_fbank_lhotsecv22_yue_captionai-generated_videoyue_and_zh_sentencesSlimPajama-6B_km_4_8_cos-d512voxbox_cosyvoice2common_voice_21_0_yuecantonese only
