datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.Movie101
Movie101
[!NOTE]
Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2
Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie.
The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.yue-logiqa
Dataset Card for Cantonese LogiQA
This dataset is a Cantonese translation of jiacheng-ye/logiqa-zh. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"context": "有啲廣東人唔鍾意食辣椒。所以,有啲南方人唔鍾意食辣椒",
"query":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-logiqa.forecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models"
This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.GuardReasonerTrain
GuardReasonerTrain
GuardReasonerTrain is the training data for R-SFT of GuardReasoner, as described in the paper GuardReasoner: Towards Reasoning-based LLM Safeguards.
Code: https://github.com/yueliu1999/GuardReasoner/
Usage
from datasets import load_dataset
# Login using e.g. `huggingface-cli login` to access this dataset
ds = load_dataset("yueliu1999/GuardReasonerTrain")
Citation
If you use this dataset, please cite our paper.
@article{GuardReasoner… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasonerTrain.MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.Yue-Benchmark
How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models
Homepage: https://github.com/jiangjyjy/Yue-Benchmark
Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark
Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models.
Introduction
The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.forecastingDataset from "Approaching Human-Level Forecasting with Language Models"
This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.aops_part1voxbox_cosyvoice2agent-outputGuardReasoner-VLTrain
GuardReasoner-VLTrain
GuardReasoner-VLTrain is the training data for R-SFT of GuardReasoner-VL, as described in the paper GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning.
Code: https://github.com/yueliu1999/GuardReasoner-VL/
Usage
from datasets import load_dataset
# Login using e.g. `huggingface-cli login` to access this dataset
ds = load_dataset("yueliu1999/GuardReasoner-VLTrain")
Citation
If you use this dataset, please cite our paper.… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/GuardReasoner-VLTrain.MeetAll
MeetAll: Bilingual Enterprise Meeting QA Dataset
Dataset Description
MeetAll is a bilingual (Chinese/English) enterprise meeting question-answering dataset from the AAAI 2026 paper "MeetBench-XL: A Benchmark for Multi-Meeting Intelligence". It contains complex QA pairs grounded in real meeting transcripts, covering 13 complexity classes across 4 dimensions.
Key Statistics
Metric
Paper Target
This Release
Total QA pairs
1,180
381
Total meetings
231… See the full description on the dataset page: https://huggingface.co/datasets/YueLinHu/MeetAll.speech2speech_vocalnetemilia_lhotse_manifestyue_xstory_cloze
Dataset Card for Cantonese XStoryCloze
This dataset is a Cantonese translation of the Simplified Chinese subset of juletxara/xstory_cloze. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"input_sentence_1":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_xstory_cloze.yue_school_math_0.25M
Cantonese School Math 0.25M
This dataset is Cantonese translation of the Simplified Chinese dataset BelleGroup/school_math_0.25M, please check the original dataset for more information.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_school_math_0.25M.yue-openrice-review
Openrice Review Classification dataset
From github.com/Christainx/Dataset_Cantonese_Openrice.
The dataset includes 60k instances from Cantonese reviews in Openrice.
The rating ranks from 1-star (very negative) to 5-star (very positive). The instances are shuffled in order to disperse reviews of same restaurant.
Code for the splits creation
import datasets
def load_openrice():
#… See the full description on the dataset page: https://huggingface.co/datasets/izhx/yue-openrice-review.FlipGuardData
FlipGuardData
This dataset contains the attack samples presented in the paper FlipAttack: Jailbreak LLMs via Flipping.
FlipAttack is a simple yet effective jailbreak attack against black-box LLMs that exploits their autoregressive nature by disguising harmful prompts using flipping transformations. FlipGuardData contains 45,000 attack samples generated against 8 different LLMs, including GPT-4o, Claude 3.5 Sonnet, and Llama 3.1.
Paper: https://huggingface.co/papers/2410.02832… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/FlipGuardData.belle_platypus_shargpt4FoodReasonSeg
FoodReasonSeg
FoodReasonSeg is built upon the food segmentation dataset FoodSeg103.
We provide the ingredients list to GPT-4 and request it to generate multi-round conversations where the questions require complex reasoning.
The corresponding masks from the original dataset for the ingredients mentioned in the answers are used as segmentation labels.
For more details, please refer to our paper FoodLMM.
Uses
Download the original FoodSeg103 dataset,
the… See the full description on the dataset page: https://huggingface.co/datasets/Yueha0/FoodReasonSeg.yue-lihkg-topicFrom https://github.com/toastynews/lihkg-cat-v2
lihkg-cat-v2
Scraped forum threads from LIHKG for categorization task. Formatted to use with BERT. Compared to v1, the number of categories increased from 18 to 20, and the number of training examples increased from 300 to 500. The minimum length for each example has also increased to make the task more solvable.
subsetzh-wiki-yue-long
Dataset Description
This dataset, named zh-wiki-yue-long, is crawled from the Yue (Cantonese) version of Wikipedia. It contains a collection of articles with an emphasis on long sentences, providing a rich source for understanding complex structures in Yue text. The dataset is designed for research in natural language processing (NLP) and machine learning tasks involving Yue text.
Data Content
Language: Yue (Cantonese)
Source: https://zh-yue.wikipedia.org
Type: Crawled… See the full description on the dataset page: https://huggingface.co/datasets/R5dwMg/zh-wiki-yue-long.answering_botlhotse_issue_1478e-girlwikitext-zh-yuereasoningweb-security
Web Security
Description
Web Security Dataset
Format
This dataset is in alpaca format.
Creation Method
This dataset was created using the Easy Dataset tool.
Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for Large Language Models (LLMs). It offers an intuitive interface for uploading domain-specific files, intelligently splitting content, generating questions, and producing high-quality training… See the full description on the dataset page: https://huggingface.co/datasets/yuej10j1anke/web-security.
