datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-zh
Dataset Card for "alpaca-zh"
本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。
Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.say-idc-media-vaultSTUZero-Atari-Dynamics
STUZero Atari Dynamics Dataset
Offline dynamics training datasets collected from trained EfficientZero V2 (EZv2) benchmark models on Atari games. Each game's data is stored in a subfolder named {game}_{steps} indicating the game and the number of training steps of the source checkpoint. While all models were trained for 120K steps, best results in some games were attained at earlier checkpoints. The model with best eval scores was used to curate data for each game.… See the full description on the dataset page: https://huggingface.co/datasets/Shivamkak/STUZero-Atari-Dynamics.cleanvid-15m_map
CleanVid Map (15M) 🎥
TempoFunk Video Generation Project
CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as:
Textual Descriptions 📃
Recording Equipment 📹
Categories 🔠
Framerate 🎞️
Aspect Ratio 📺
CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it.
This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.ZDPShift
ZDPShift: Beyond the Zero-Disparity Plane in Stereo
Every public stereo benchmark assumes positive disparity valuesd = fB/Z ≥ 0. Mordern stereoscopic display — cinema 3D, VR, HMDs — actively uses d < 0. ZDPShift bridges the gap: the same artist-authored open-movie content rendered at five ZDP shifts Δ ∈ {−16, 0, +16, +24, +32} pixels, giving you a controlled continuum from textbook-positive to substantially-crossed disparities, with analytical ground truth at every pixel.… See the full description on the dataset page: https://huggingface.co/datasets/shijianjian/ZDPShift.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.shiunjikenokodomotachi
Bangumi Image Base of Shiunji-ke No Kodomotachi
This is the image base of bangumi Shiunji-ke no Kodomotachi, we detected 39 characters, 4423 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiunjikenokodomotachi.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.gpt-4v-distribution-shift
License
This repository is licensed under the MIT License.
Description
This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift.
These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios.
Using the Dataset
For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.bge-m3-data
Dataset Summary
This depository contains all the fine-tuning data for the bge-m3 model, including:
Dataset
Language
MS MARCO
English
NQ
English
HotpotQA
English
TriviaQA
English
SQuAD
English
COLIEE
English
PubMedQA
English
NLI from SimCSE
English
DuReader
Chinese
mMARCO-zh
Chinese
T2Ranking
Chinese
Law-GPT
Chinese
cMedQAv2
Chinese
NLI-zh
Chinese
LeCaRDv2
Chinese
Mr.TyDi
11 languages
MIRACL
16 languages
MLDR
13 languages
Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.aitod-v2
AI-TOD-v2
AI-TOD-v2, the tiny-object detection benchmark in aerial images, packed once with the official v2 annotations kept whole, so it loads in one line and no data path has to be configured:
from datasets import load_dataset
ds = load_dataset("shijli/aitod-v2") # 11214 train / 2804 validation / 14018 test
AI-TOD cuts 28036 images of 800 x 800 pixels from xView, DOTA-v1.5, VisDrone2018-Det, Airbus Ship Detection and DIOR, and annotates eight classes whose mean object size… See the full description on the dataset page: https://huggingface.co/datasets/shijli/aitod-v2.shiftproject_test_beirBEIR version of vidore/shiftproject_test.
nli_zh纯文本数据,格式:(sentence1, sentence2, label)。常见中文语义匹配数据集,包含ATEC、BQ、LCQMC、PAWSX、STS-B共5个任务。VLM4D
VLM4D
VLM4D is a benchmark for evaluating the spatiotemporal reasoning capabilities of Vision Language Models (VLMs). It contains real and synthetic videos paired with multiple-choice questions that require models to reason about translation, rotation, perspective, motion continuity, counting, and false-positive events.
The dataset was introduced in VLM4D: Towards Spatiotemporal Awareness in Vision Language Models, accepted to ICCV 2025.
Project page: https://vlm4d.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/shijiezhou/VLM4D.tuxun
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
GeoComp
Dataset description
Inspired by geoguessr.com, we developed a free geolocation game platform that tracks participants' competition histories.
Unlike most geolocation websites, including Geoguessr, which rely solely on samples from Google Street… See the full description on the dataset page: https://huggingface.co/datasets/ShirohAO/tuxun.physical-ai-bench-conditional-generation
Physical AI Bench - Conditional Generation
Paper | Code
This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.
Dataset
Category
Sample Nums
Agibot World
Robotics
200
OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.shijousaikyounodaimaoumurabitoanitenseisuru
Bangumi Image Base of Shijou Saikyou No Daimaou, Murabito A Ni Tensei Suru
This is the image base of bangumi Shijou Saikyou no Daimaou, Murabito A ni Tensei suru, we detected 70 characters, 4747 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shijousaikyounodaimaoumurabitoanitenseisuru.skin-cancer-ham10000-datasetTCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.nli-zh-allThe SNLI corpus (version 1.0) is a merged chinese sentence similarity dataset, supporting the task of natural language
inference (NLI), also known as recognizing textual entailment (RTE).shiguangdailirenii
Bangumi Image Base of Shiguang Dailiren Ii
This is the image base of bangumi Shiguang Dailiren II, we detected 48 characters, 3677 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiguangdailirenii.W_LSTMix_test_datasetshirobako
Bangumi Image Base of Shirobako
This is the image base of bangumi Shirobako, we detected 52 characters, 3771 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shirobako.mmconflict-editable-values-1k
MMConflict Editable Values 2K
This dataset contains 2,000 source images with visible atomic values for
multimodal conflict research. It has 100 images in each of 20 categories. Every
image comes from a photograph, scan, captured website, software screenshot, or
page of a source document. The dataset does not contain generated images or
project-rendered examples.
Each row records the source, source URL, license, attribution, visible value,
question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.shinnonakamajanaitoyuushanopartywooidasaretanodehenkyoudeslowlifesurukotonishimashita2nd
Bangumi Image Base of Shin No Nakama Ja Nai To Yuusha No Party Wo Oidasareta Node, Henkyou De Slow Life Suru Koto Ni Shimashita 2nd
This is the image base of bangumi Shin no Nakama ja Nai to Yuusha no Party wo Oidasareta node, Henkyou de Slow Life suru Koto ni Shimashita 2nd, we detected 69 characters, 4925 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shinnonakamajanaitoyuushanopartywooidasaretanodehenkyoudeslowlifesurukotonishimashita2nd.shiftproject_test_beirBEIR version of vidore/shiftproject_test.
MathCanvas-Edit
MathCanvas-Edit Dataset
🚀 Data Usage
from datasets import load_dataset
dataset = load_dataset("shiwk24/MathCanvas-Edit")
print(dataset)
📖 Overview
MathCanvas-Edit is a large-scale dataset containing 5.2 million step-by-step editing trajectories, forming a crucial component of the [MathCanvas] framework. MathCanvas is designed to endow Unified Large Multimodal Models (LMMs) with intrinsic… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Edit.MathCanvas-Instruct
MathCanvas-Instruct Dataset
🚀 Data Usage
from datasets import load_dataset
dataset = load_dataset("shiwk24/MathCanvas-Instruct")
print(dataset)
📖 Overview
MathCanvas-Instruct is a high-quality, fine-tuning dataset with 219K examples of interleaved visual-textual reasoning paths. It is the core component for the second phase of the [MathCanvas] framework: Strategic Visual-Aided Reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Instruct.
