CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes5.4k downloads19d agoHugging Face02yuezih /Movie101gated Movie101 [!NOTE] Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2 Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie. The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.imagevideo-text-to-text100K<n<1M7 likes4.4k downloads1y agoHugging Face03ming030890 /youtube_caption_yue YouTube ASR Caption Dataset (Cantonese) This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs. What’s included Segments where the ASR output is identical to the original caption — likely clean. Segments where differences are only homophones (同音字) or English words — likely ASR mistakes. This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.audio10K<n<100K2 likes3.6k downloads1y agoHugging Face04Yuehao /SlimPajama-6B_km-ip-d512tabular1M<n<10M0 likes727 downloads9mo agoHugging Face05Yuehao /celeba CelebA dataset A copy of celeba dataset. https://mmlab.ie.cuhk.edu.hk/projects/CelebA.html How to use Download data huggingface-cli download --local-dir /path/to/datasets/celeba --repo-type dataset Yuehao/celeba unzip /path/to/datasets/celeba/img_align_celeba.zip -d /path/to/datasets/celeba Load data via torchvision.datasets.CelebA torchvision.datasets.CelebA(root='/path/to/datasets') image100K<n<1M0 likes589 downloads1y agoHugging Face06JackyHoCL /common_voice_22_yue2025-08-03 Update: Use MP3 instead of WAV All Right reserved by mozilla-foundation audio100K<n<1M1 likes491 downloads1y agoHugging Face07hon9kon9ize /yue_emo_speech Cantonese Emotional Speech Crawled from YouTube and RTHK, this dataset contains 1,000 hours of Cantonese speech, each labeled with one of the following emotions: angry, disgusted, fearful, happy, neutral, other, sad, or surprised. The dataset also includes the confidence of the emotion label. The audio files are denoised with resemble-enhance. The transcriptions are generated by SenseVoiceSmall, and deduplicated using MinHash. audio100K<n<1M16 likes416 downloads2y agoHugging Face08yuekai /speechio_test SpeechIO ASR Test Sets (parquet) Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark, re-exported from yuekai/speechio (Lhotse cuts) into standard HuggingFace parquet with embedded 16 kHz audio. 27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split. ~43k utterances, ~66 hours total, evaluation only. Columns column type note segment_id string utterance id speaker string speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.audioautomatic-speech-recognition10K<n<100K0 likes406 downloads3mo agoHugging Face09yuekai /InstructS2S-200Kaudio100K<n<1M1 likes397 downloads1y agoHugging Face10hon9kon9ize /yue-logiqa Dataset Card for Cantonese LogiQA This dataset is a Cantonese translation of jiacheng-ye/logiqa-zh. For more detailed information about the original dataset, please refer to the provided link. This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset. Sample { "context": "有啲廣東人唔鍾意食辣椒。所以,有啲南方人唔鍾意食辣椒", "query":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-logiqa.text1K<n<10K1 likes384 downloads3y agoHugging Face11YuehHanChen /forecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models" This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. Data Source and Format The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.text10K<n<100K7 likes355 downloads3y agoHugging Face12yuekai /multi_en_qwen3_omni_sglang_regeneratedaudio100K<n<1M0 likes346 downloads2mo agoHugging Face13yuecao0119 /MMInstruct-GPT4V MMInstruct The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity". The data engine is available on GitHub at yuecao0119/MMInstruct. Todo List Data Engine. Open Source Datasets. Release the checkpoint. Introduction Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.imagevisual-question-answering100K<n<1M13 likes334 downloads2y agoHugging Face14yueliu1999 /GuardReasonerTrain GuardReasonerTrain GuardReasonerTrain is the training data for R-SFT of GuardReasoner, as described in the paper GuardReasoner: Towards Reasoning-based LLM Safeguards. Code: https://github.com/yueliu1999/GuardReasoner/ Usage from datasets import load_dataset # Login using e.g. `huggingface-cli login` to access this dataset ds = load_dataset("yueliu1999/GuardReasonerTrain") Citation If you use this dataset, please cite our paper. @article{GuardReasoner… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasonerTrain.texttext-classification100K<n<1M4 likes324 downloads1y agoHugging Face15jiuyi111 /YuE2-Musictextn<1K1 likes317 downloads4d agoHugging Face16BillBao /Yue-Benchmark How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models Homepage: https://github.com/jiangjyjy/Yue-Benchmark Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models. Introduction The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.textmultiple-choice1K<n<10K8 likes313 downloads2y agoHugging Face17yuekai /CV3-Evalaudio1K<n<10K2 likes311 downloads1y agoHugging Face18JackyHoCL /common_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps TODO: convert to MP3, reduce size audio100K<n<1M0 likes275 downloads5mo agoHugging Face19YueHuLab /AlphaPanda_checkpointstext100K<n<1M0 likes271 downloads2y agoHugging Face20YuehHanChen /forecastingDataset from "Approaching Human-Level Forecasting with Language Models" This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. Data Source and Format The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.text1K<n<10K12 likes216 downloads3y agoHugging Face21yuekai /seed_tts_cosy2audio1K<n<10K1 likes203 downloads2y agoHugging Face22yuerxin /aops_part1text1K<n<10K0 likes180 downloads6mo agoHugging Face23J017athan /Multimodal-Yue-Benchmark Multimodal Yue Benchmark Cantonese audio + text benchmark derived from BillBao/Yue-Benchmark (Yue-GSM8K & Yue-MMLU-style tasks). We kept the same task content in Cantonese and added TTS for three Cantonese speakers (hiugaai, hiumaan, wanlung). Subsets & splits Config name Task Speaker Splits mmlu_* multiple-choice (MMLU-style) per speaker train, test gsm8k_* math word problems (GSM8K-style) per speaker train, test Example: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/J017athan/Multimodal-Yue-Benchmark.audio10K<n<100K0 likes152 downloads6mo agoHugging Face24yuekai /alimeeting_aishell4_training_whisper_fbank_lhotseimagen<1K1 likes143 downloads2y agoHugging Face25JackyHoCL /cv22_yue_captionaudio100K<n<1M0 likes131 downloads6mo agoHugging Face26yueying-117 /ai-generated_videotextn<1K0 likes125 downloads11mo agoHugging Face27AlienKevin /yue_and_zh_sentencestext10M<n<100M1 likes120 downloads2y agoHugging Face28Yuehao /SlimPajama-6B_km_4_8_cos-d512tabular1M<n<10M0 likes117 downloads8mo agoHugging Face29yuekai /voxbox_cosyvoice2tabular10M<n<100M1 likes108 downloads1y agoHugging Face30ming030890 /common_voice_21_0_yuecantonese only audio10K<n<100K0 likes107 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.