datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.Yue-Benchmark
How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models
Homepage: https://github.com/jiangjyjy/Yue-Benchmark
Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark
Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models.
Introduction
The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.MLLM-CITBench
MLLM-CITBench Multimodal Task Benchmarking Dataset
This dataset contains 7 tasks:
OCR: Optical Character Recognition task
art: Art - related task
fomc: Financial and Monetary Policy - related task
math: Mathematical problem - solving task
medical: Medical - related task
numglue: Numerical reasoning task
science: Scientific problem - solving task
Each task has independent training and test splits. The image data is stored in the dataset in the form of full file.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/yueluoshuangtian/MLLM-CITBench.MeetAll
MeetAll: Bilingual Enterprise Meeting QA Dataset
Dataset Description
MeetAll is a bilingual (Chinese/English) enterprise meeting question-answering dataset from the AAAI 2026 paper "MeetBench-XL: A Benchmark for Multi-Meeting Intelligence". It contains complex QA pairs grounded in real meeting transcripts, covering 13 complexity classes across 4 dimensions.
Key Statistics
Metric
Paper Target
This Release
Total QA pairs
1,180
381
Total meetings
231… See the full description on the dataset page: https://huggingface.co/datasets/YueLinHu/MeetAll.MeetAll-v2
MeetAll-v2: A Repaired Reproduction of the MeetAll Benchmark
Version Notice
This is MeetAll-v2, a repaired/reproduced version of the MeetAll cross-meeting understanding benchmark. It is NOT an exact byte-for-byte copy of the original paper release. See Reproduction Status for details.
Dataset Description
MeetAll is a benchmark for evaluating cross-meeting understanding capabilities of language models. It is built on top of real meeting transcripts… See the full description on the dataset page: https://huggingface.co/datasets/YueLinHu/MeetAll-v2.EmoSupportBench
EmoSupportBench
EmoSupportBench is a comprehensive dataset and benchmark for evaluating emotional support capabilities of large language models (LLMs). It provides a systematic framework to assess how well AI systems can provide empathetic, helpful, and psychologically-grounded support to users seeking emotional assistance.
🎯 Key Features
200-question bilingual evaluation set (English & Chinese) covering 8 major emotional support scenarios
Hierarchical scenario… See the full description on the dataset page: https://huggingface.co/datasets/YueyangWang/EmoSupportBench.MedPAIR
Dataset Card for MedPAIR
MedPAIR represents a "Medical Dataset Comparing Physician Trainees and AI Relevance Estimation and Question Answering". We design MedPAIR to compare LLM reasoning processes to those of physician trainees and to enable future research to focus on relevant features. MedPAIR is a first benchmark step to matching the relevancy annotated by clinical professional labelers to that estimated by LLMs. The motivation for MedPAIR is to ensure that what the LLM finds… See the full description on the dataset page: https://huggingface.co/datasets/YuexingHao/MedPAIR.pittsburgh_floods_street_levelThe dataset provides fine-grained spatiotemporal information on urban floods occurring inside the city of Pittsburgh, PA, USA, from 2015 to 2024 by integrating publicly available data sources. The data sources include NOAA storm events database and Pittsburgh 311 flooding requests. Each row corresponds to one segment flood event, characterized by the street segment defined by a distinct combination of "u_node", "v_node", "length_m", and time. Each flood event was mapped to the street segments… See the full description on the dataset page: https://huggingface.co/datasets/yueq92/pittsburgh_floods_street_level.rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for.
Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py.
