CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes5.8k downloads21d agoHugging Face02yuezih /Movie101gated Movie101 [!NOTE] Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2 Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie. The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.imagevideo-text-to-text100K<n<1M7 likes4.5k downloads1y agoHugging Face03hon9kon9ize /yue-logiqa Dataset Card for Cantonese LogiQA This dataset is a Cantonese translation of jiacheng-ye/logiqa-zh. For more detailed information about the original dataset, please refer to the provided link. This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset. Sample { "context": "有啲廣東人唔鍾意食辣椒。所以,有啲南方人唔鍾意食辣椒", "query":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-logiqa.text1K<n<10K1 likes385 downloads3y agoHugging Face04YuehHanChen /forecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models" This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. Data Source and Format The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.text10K<n<100K7 likes371 downloads3y agoHugging Face05yueliu1999 /GuardReasonerTrain GuardReasonerTrain GuardReasonerTrain is the training data for R-SFT of GuardReasoner, as described in the paper GuardReasoner: Towards Reasoning-based LLM Safeguards. Code: https://github.com/yueliu1999/GuardReasoner/ Usage from datasets import load_dataset # Login using e.g. `huggingface-cli login` to access this dataset ds = load_dataset("yueliu1999/GuardReasonerTrain") Citation If you use this dataset, please cite our paper. @article{GuardReasoner… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasonerTrain.texttext-classification100K<n<1M4 likes342 downloads1y agoHugging Face06yuecao0119 /MMInstruct-GPT4V MMInstruct The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity". The data engine is available on GitHub at yuecao0119/MMInstruct. Todo List Data Engine. Open Source Datasets. Release the checkpoint. Introduction Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.imagevisual-question-answering100K<n<1M13 likes333 downloads2y agoHugging Face07BillBao /Yue-Benchmark How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models Homepage: https://github.com/jiangjyjy/Yue-Benchmark Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models. Introduction The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.textmultiple-choice1K<n<10K8 likes316 downloads2y agoHugging Face08YuehHanChen /forecastingDataset from "Approaching Human-Level Forecasting with Language Models" This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. Data Source and Format The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.text1K<n<10K12 likes221 downloads3y agoHugging Face09yuerxin /aops_part1text1K<n<10K0 likes184 downloads6mo agoHugging Face10yuekai /voxbox_cosyvoice2tabular10M<n<100M1 likes110 downloads1y agoHugging Face11yueqis /agent-outputtext1K<n<10K0 likes81 downloads2y agoHugging Face12yueliu1999 /GuardReasoner-VLTrain GuardReasoner-VLTrain GuardReasoner-VLTrain is the training data for R-SFT of GuardReasoner-VL, as described in the paper GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning. Code: https://github.com/yueliu1999/GuardReasoner-VL/ Usage from datasets import load_dataset # Login using e.g. `huggingface-cli login` to access this dataset ds = load_dataset("yueliu1999/GuardReasoner-VLTrain") Citation If you use this dataset, please cite our paper.… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasoner-VLTrain.textimage-text-to-text100K<n<1M1 likes70 downloads1y agoHugging Face13YueLinHu /MeetAll MeetAll: Bilingual Enterprise Meeting QA Dataset Dataset Description MeetAll is a bilingual (Chinese/English) enterprise meeting question-answering dataset from the AAAI 2026 paper "MeetBench-XL: A Benchmark for Multi-Meeting Intelligence". It contains complex QA pairs grounded in real meeting transcripts, covering 13 complexity classes across 4 dimensions. Key Statistics Metric Paper Target This Release Total QA pairs 1,180 381 Total meetings 231… See the full description on the dataset page: https://huggingface.co/datasets/YueLinHu/MeetAll.textquestion-answeringn<1K0 likes70 downloads5mo agoHugging Face14yuekai /speech2speech_vocalnettabular1K<n<10K0 likes66 downloads1y agoHugging Face15yuekai /emilia_lhotse_manifesttabular10M<n<100M1 likes59 downloads2y agoHugging Face16hon9kon9ize /yue_xstory_cloze Dataset Card for Cantonese XStoryCloze This dataset is a Cantonese translation of the Simplified Chinese subset of juletxara/xstory_cloze. For more detailed information about the original dataset, please refer to the provided link. This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset. Sample { "input_sentence_1":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_xstory_cloze.text1K<n<10K1 likes49 downloads3y agoHugging Face17hon9kon9ize /yue_school_math_0.25M Cantonese School Math 0.25M This dataset is Cantonese translation of the Simplified Chinese dataset BelleGroup/school_math_0.25M, please check the original dataset for more information. This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset. Sample { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_school_math_0.25M.text100K<n<1M1 likes48 downloads3y agoHugging Face18izhx /yue-openrice-review Openrice Review Classification dataset From github.com/Christainx/Dataset_Cantonese_Openrice. The dataset includes 60k instances from Cantonese reviews in Openrice. The rating ranks from 1-star (very negative) to 5-star (very positive). The instances are shuffled in order to disperse reviews of same restaurant. Code for the splits creation import datasets def load_openrice(): #… See the full description on the dataset page: https://huggingface.co/datasets/izhx/yue-openrice-review.text10K<n<100K1 likes47 downloads2y agoHugging Face19yueliu1999 /FlipGuardData FlipGuardData This dataset contains the attack samples presented in the paper FlipAttack: Jailbreak LLMs via Flipping. FlipAttack is a simple yet effective jailbreak attack against black-box LLMs that exploits their autoregressive nature by disguising harmful prompts using flipping transformations. FlipGuardData contains 45,000 attack samples generated against 8 different LLMs, including GPT-4o, Claude 3.5 Sonnet, and Llama 3.1. Paper: https://huggingface.co/papers/2410.02832… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/FlipGuardData.tabulartext-generation10K<n<100K3 likes42 downloads4mo agoHugging Face20yuekai /belle_platypus_shargpt4text10K<n<100K3 likes32 downloads3y agoHugging Face21Yueha0 /FoodReasonSeg FoodReasonSeg FoodReasonSeg is built upon the food segmentation dataset FoodSeg103. We provide the ingredients list to GPT-4 and request it to generate multi-round conversations where the questions require complex reasoning. The corresponding masks from the original dataset for the ingredients mentioned in the answers are used as segmentation labels. For more details, please refer to our paper FoodLMM. Uses Download the original FoodSeg103 dataset, the… See the full description on the dataset page: https://huggingface.co/datasets/Yueha0/FoodReasonSeg.text1K<n<10K1 likes28 downloads2y agoHugging Face22izhx /yue-lihkg-topicFrom https://github.com/toastynews/lihkg-cat-v2 lihkg-cat-v2 Scraped forum threads from LIHKG for categorization task. Formatted to use with BERT. Compared to v1, the number of categories increased from 18 to 20, and the number of training examples increased from 300 to 500. The minimum length for each example has also increased to make the task more solvable. text10K<n<100K0 likes27 downloads2y agoHugging Face23yueqis /subsettextn<1K0 likes26 downloads1y agoHugging Face24R5dwMg /zh-wiki-yue-long Dataset Description This dataset, named zh-wiki-yue-long, is crawled from the Yue (Cantonese) version of Wikipedia. It contains a collection of articles with an emphasis on long sentences, providing a rich source for understanding complex structures in Yue text. The dataset is designed for research in natural language processing (NLP) and machine learning tasks involving Yue text. Data Content Language: Yue (Cantonese) Source: https://zh-yue.wikipedia.org Type: Crawled… See the full description on the dataset page: https://huggingface.co/datasets/R5dwMg/zh-wiki-yue-long.text10K<n<100K2 likes25 downloads2y agoHugging Face25yueyulin /answering_bottext100K<n<1M1 likes25 downloads2y agoHugging Face26yuekai /lhotse_issue_1478tabularn<1K0 likes25 downloads1y agoHugging Face27yuebanlaosiji /e-girltextn<1K0 likes21 downloads2y agoHugging Face28indiejoseph /wikitext-zh-yuetext100K<n<1M1 likes19 downloads3y agoHugging Face29yueqis /reasoningtext1K<n<10K0 likes19 downloads2y agoHugging Face30yuej10j1anke /web-security Web Security Description Web Security Dataset Format This dataset is in alpaca format. Creation Method This dataset was created using the Easy Dataset tool. Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for Large Language Models (LLMs). It offers an intuitive interface for uploading domain-specific files, intelligently splitting content, generating questions, and producing high-quality training… See the full description on the dataset page: https://huggingface.co/datasets/yuej10j1anke/web-security.text1K<n<10K1 likes19 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.